Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
TensorX logo

Model details

Kimi K3

Kimi K3 is a frontier-scale sparse Mixture-of-Experts model built by Moonshot AI, designed for long-horizon agentic work rather than casual chat. At roughly 2.8 trillion total parameters with only 16 of 896 experts active per token, it relies on a hybrid attention stack that interleaves Kimi Delta Attention with global attention layers and uses Attention Residuals to keep information flowing across depth. The architectural focus on routing balance and linear attention is explicitly aimed at making a million-token context window practical at training and inference time, and independent reporting describes the model as competitive with leading closed systems on coding, tool-use, and document-heavy benchmarks while leaving room for stronger general reasoning results in future iterations.

Practically, Kimi K3 fits workloads that span large codebases, long research sessions, multimodal evidence review, and sustained tool-driven automation, where its native vision input, generous context, and published checkpoint give teams flexibility to inspect, fine-tune, or self-deploy. The published weights, technical report, and serving recipes reflect an emphasis on openness and architectural transparency, though deployment still calls for significant accelerator infrastructure and pricing sits well above smaller MoE alternatives, so it is best treated as a specialist escalation model for the hardest coding, research, and visual-agent tasks rather than a default for routine traffic.

TensorXmoonshotai/kimi-k3kimi-k3

Quick Info

Powered by
Provider
TensorX
Model key
moonshotai/kimi-k3
Release date
Jul 16, 2026
Last updated
Jul 16, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$3.00
Output token cost
$15.00

Limits

Output tokens
131,072 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare Kimi K3 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Kimi K3

GreenPT

Coverage

The Moonshot-authored technical report mirrored on alphaXiv explicitly names Kimi K3 and documents the exact model specifications: a 2.8 trillion parameter Mixture-of-Experts architecture with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. The report attributes devel Post-training methodology encompasses reinforcement learning across general, agentic, and coding domains with multiple reasoning-effort levels, enabling compositional generalization and robust long-horizon execution. The paper notes infrastructure advances including algorithm-system co-design for KDA, perfectly balance

TensorX

CoverageBenchmark

Trilogy AI's newsletter confirms Kimi K3 is live through Moonshot's API, kimi.com, the Kimi mobile apps, Kimi Work, Kimi Code, and OpenRouter, giving developers multiple access paths including for those without a Moonshot account. Independent evaluation places K3 at an Artificial Analysis Intelligence Index score of 57 Trilogy AI's coverage reaffirms the core specifications: 2.8 trillion parameters, native vision, a 1,048,576-token context window, $3/$15 API pricing, and a Stable LatentMoE with 896 experts activating 16 per token. The launch is described as a long-horizon agent model for software engineering, knowledge work, and mult

Alibaba (China)

Coverage

The Geopolitechs post documents Kimi K3 Open Day on July 27, 2026, when Moonshot released the model weights, technical report, and supporting infrastructure technologies MoonEP, FlashKDA, and AgentEnv. The excerpt confirms K3 as a 2.8-trillion-parameter Mixture-of-Experts model with native visual understanding and a 1- The technical report excerpts describe KDA mixed with Gated MLA at a 3:1 ratio for efficient long-context modeling with block-level attention residuals improving cross-layer information flow. The Stable LatentMoE architecture activates 16 out of 896 routed experts per token, using SiTU-GLU and Quantile Balancing to mai

Vancine

Coverage

ExplainX.ai reported that Moonshot AI released free, public Kimi K3 weights on July 26, 2026, at roughly 7:30 PM EDT, a day ahead of the previously communicated July 27 target. Confirmed specs drawn from Moonshot's Kimi K3 platform pages include 2.8 trillion parameters and a 1,048,576-token (1M) context window, followi The explainer covers an architecture deep-dive from Sebastian Raschka covering LatentMoE, NoPE, KDA, and attention residuals, plus a laptop existence-proof demonstration of Deltafin running on an M1 Max at 16 seconds per token. It also notes Anthropic CEO Dario Amodei's July 27 statement that Anthropic has never advoca

Videos about Kimi K3

More models around Kimi K3