Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Wallaby logo

Model details

Kimi K3

Kimi K3 is designed as an open-weight, native multimodal model for long-running software engineering, knowledge work, and reasoning tasks. Its architecture combines Kimi Delta Attention with Attention Residuals, while a Stable LatentMoE setup activates 16 of 896 experts. The model has 2.8 trillion total parameters and native vision capabilities, making it suited to workflows that combine textual and visual information rather than treating images as an add-on.

For practical use, Kimi K3 is aimed at agents that work across substantial projects and extended context. Its native vision, million-token context, and benchmark results on software development, research, spreadsheets, browsing, and long-context tasks support a focus on complex multi-step work. The model is an open-weight release, which makes its architecture and checkpoint available for independent deployment and evaluation.

Wallabymoonshotai/kimi-k3kimi-k3

Quick Info

Powered by
Provider
Wallaby
Model key
moonshotai/kimi-k3
Release date
Jul 16, 2026
Last updated
Jul 16, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$2.70
Output token cost
$13.50

Limits

Output tokens
131,072 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare Kimi K3 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Kimi K3

OpenRouter

Coverage

On July 27, 2026, Moonshot AI held Kimi K3 Open Day, publicly releasing the full Kimi K3 model weights alongside a technical report and supporting infrastructure stack (MoonEP, FlashKDA, AgentEnv). The post confirms Kimi K3 as a 2.8-trillion-parameter Mixture-of-Experts model with native visual understanding and a 1-mi The accompanying technical report details the architecture used to support long context and high sparsity: KDA and Gated MLA are mixed at a 3:1 ratio for efficient long-context modeling, with block-level attention residuals improving cross-layer information flow. Each token activates 16 of 896 routed experts, with SiTU

OpenRouter

CoverageBenchmark

Yotta Labs' technical breakdown confirms Kimi K3 as Moonshot AI's successor to the K2 line, announced mid-July 2026 with API access first and open weights shipped July 27, 2026 as scheduled. Core specifications are 2.8 trillion total parameters in a sparse Mixture-of-Experts layout with 896 experts (16 active per token Architecturally, Kimi K3 introduces a hybrid linear attention design called Kimi Delta Attention aimed at making the 1M-token context affordable to serve, and Moonshot claims a 2.5× scaling-efficiency improvement over K2. The Kimi K3 License permits free commercial use with attribution, with extra conditions only at ve

OpenRouter

CoverageBenchmark

BenchmarkList's aggregator page shows Kimi K3 at $3 input / $15 output per million tokens, with 124 benchmark rows. On GDPval-AA, Kimi K3 scores an Elo of 1668 in max-effort mode (2nd of 34, behind Claude Opus 5 at 1861) and 1271 in low-effort mode. Tau3-Banking records 46.0% Pass@1 (8th of 17, field leader GLM 5.3 at The page compares Kimi K3 against Claude Fable 5.1, Claude Opus 5, Qwen3.8-Flash-Next, Qwen3.8-2.4T-A95B, GLM 5.3, GLM 5.3 Flash, and DeepSeek V4 Flash 0731, with benchmark entries dated through September 2, 2026. Eval data is sourced from Artificial Analysis. The page documents that most detailed scores use vendor-rep

OpenRouter

Coverage

Nathan Lambert's Interconnects analysis (July 20, 2026) framed Moonshot AI's July 16 Kimi K3 release as a true frontier model — the closest open models have been to the frontier since DeepSeek R1 — and called it "clearly the strongest open model ever released." K3 placed #2 on the Vals AI index, #3 on Artificial Analys K3 is a 2.8-trillion-parameter mixture-of-experts model whose open-weight release was scheduled for July 27, and Lambert conditioned his ecosystem-level analysis on Moonshot honoring that commitment. He characterized K3 as an example of a Chinese lab executing on scaling known axes — data, algorithms, architecture, too

OpenRouter

CoverageBenchmark

India Today's July 17, 2026 report confirms Moonshot AI's launch of Kimi K3, which the company describes as the largest open-weight AI model at 2.8 trillion parameters. Moonshot positions Kimi K3 as a substantial improvement over prior versions, claiming it can match or outperform the most advanced US models including The article explains that "open-weight" means the Kimi K3 model can be freely downloaded and run locally, in contrast to closed-weight competitors like Fable and GPT-5.6 that are cloud-only and subject to token-based pricing. The 2.8 trillion parameter scale is highlighted as Kimi K3's differentiating factor relative t

Videos about Kimi K3

More models around Kimi K3