Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
LLM Gateway logo

Model details

Kimi K3 (Runpod)

Kimi K3 is the latest open-source model from Moonshot Labs, notable for combining an extreme Mixture-of-Experts scale with a hybrid attention design. The model carries roughly 2.8 trillion parameters across 93 layers and is described as pushing the frontier of sparse architectures paired with linear attention, signaling a shift in how very large open-weight models can be organized. With released weights, it continues Moonshot's track record of setting the upper bound for openly available model sizes, narrowing the practical gap between closed APIs and self-hosted deployments for organizations that want full control over their stack.

In practical terms, Kimi K3 is positioned for advanced coding, knowledge work, and reasoning workloads, and it brings a one-million-token context window that supports long-document analysis, multi-file codebases, and extended multi-turn conversations without losing earlier content. The combination of extreme sparsity and a hybrid linear attention stack is intended to make such long contexts tractable at inference time while preserving reasoning depth. For teams looking to self-host frontier-class reasoning and coding capabilities without relying on closed APIs, Kimi K3 offers a compelling open-weight foundation that can be deployed and audited on their own infrastructure.

LLM Gatewayrunpod/kimi-k3kimi-k3

Quick Info

Powered by
Provider
LLM Gateway
Model key
runpod/kimi-k3
Release date
Jul 16, 2026
Last updated
Jul 16, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$3.00
Output token cost
$15.00

Limits

Output tokens
1,048,576 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare Kimi K3 (Runpod) pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Kimi K3 (Runpod)

LLM Gateway

CoverageBenchmark

The vLLM project blog (September 13, 2026) documents a major performance pass for serving Kimi K3, reporting 56%–60% lower latency, 2.2–2.8× higher throughput, and 72%–85% lower TTFT versus vLLM v0.27.1 on an 8xB300 node with CUDA 13.3, TP8, and an 8K/1K workload with 8-token DSpark speculative decoding across concurre Concrete optimizations are listed with PR numbers: PR 51725 introduces an adaptive speculative-token budget, PR 51726 lifts `max_num_batched_tokens` from 8,192 to 16,384 on the high-memory GPU tier (TTFT down 55–65%, throughput up 41.5%), plus internal KDA prefix checkpoints, zero-copy mixed KDA batches, deferred MXFP4

LLM Gateway

CoverageAnalysis

Marketersindex's technical analysis (August 24, 2026) frames the Kimi K3 release as a shift toward granular user control and architectural efficiency, highlighting Moonshot AI's introduction of a "reasoning effort" parameter with low, high, and max tiers. When set to "max" reasoning effort, Kimi K3 can use up to 131,07 The piece describes Kimi K3 as a 2.8-trillion-parameter Mixture-of-Experts model — among the largest publicly available — using an expert routing scheme with only 16 experts active per token, building on Moonshot's long-context lineage from earlier Kimi Chat iterations. It positions the model as a fundamental redesign

LLM Gateway

Coverage

Byteiota's technical overview describes Kimi K3 as a 2.8-trillion-parameter sparse Mixture-of-Experts model released by Moonshot AI on July 27, 2026, with 104 billion parameters active per token, 896 experts of which 16 fire per forward pass, a 1-million-token context window, and native text-and-image multimodality thr The article details two architectural additions over K2: Kimi Delta Attention (KDA) for long-context retrieval efficiency, and AttnRes working depth-wise to selectively retrieve representations, with Moonshot claiming roughly 25% better training efficiency at under 2% additional cost and a 2.5× scaling-efficiency impro

LLM Gateway

Coverage

Layer3 Labs' local-setup guide (updated September 7, 2026) states that running Kimi K3 locally is a serious multi-GPU server project — a 2.8-trillion-parameter model that cannot fit on a laptop or single consumer GPU — and emphasizes that local hosting requires holding the full weights in memory across the cluster even The guide places Kimi K3's open-weights release in the context of regulated teams that may want self-hosting, discusses hardware sizing and quantization, and positions vLLM and Ollama as relevant tools while recommending hosted API endpoints for teams without existing GPU infrastructure. It cross-references the compani

LLM Gateway

CoverageBenchmark

Layer3 Labs' benchmark guide (updated September 7, 2026) separates one independently verified Kimi K3 result from vendor-reported claims: LMArena's Frontend Code Arena first-place finish with 1,679 points is the sole third-party confirmation, a blind human-preference test specific to front-end code generation. The guid The article frames frontier-model evaluations across four categories — coding, agentic tool-use, reasoning and math, and long-context retrieval — and advises readers to consult current independent sources rather than relying on secondhand summaries. It positions K3's documented strength in front-end code generation as

LLM Gateway

Coverage

Runpod published a deployment guide confirming that Moonshot AI's Kimi K3 (released July 27, 2026 as a 2.8-trillion-parameter open-weight MoE model) can be served on a single 8xB300 pod, with about 800 GB of headroom for KV cache and context. The post establishes that 8xH100 (640 GB) and 8xB200 (1,536 GB) both fail to The guide details that Kimi K3 was quantization-aware trained with native MXFP4 weights and MXFP8 activations, so no BF16 conversion is needed, and recommends the vLLM Docker image `vllm/vllm-openai:kimi-k3` on CUDA 13 (no working pip install path because the integration depends on pre-release FlashInfer). A verified s

Videos about Kimi K3 (Runpod)

More models around Kimi K3 (Runpod)