Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
LLM Gateway logo

Model details

Kimi K3 (Tencent Cloud)

Kimi K3 is a Mixture-of-Experts model with roughly 2.8 trillion total parameters, making it the largest open-weight release from any lab to date and the anchor of Moonshot AI's push into the multi-trillion-parameter tier. The architecture activates only a fraction of those parameters per token, which keeps inference costs manageable relative to the headline scale. Open distribution of the weights was a deliberate strategy: developers and enterprises can self-host, fine-tune, and deploy the model without relying on a proprietary API, and the weights are positioned to attract ecosystem adoption across the Chinese open-model community.

Kimi K3 was designed with long-context work in mind, supporting context windows up to one million tokens and native vision, which puts it in the same bracket as other frontier long-context systems aimed at large codebases, document analysis, and multi-turn agentic workflows. In independent head-to-head evaluations, the model has shown competitive results against leading Western frontier systems and has been selected by Western developer-tooling companies as a foundation for production coding assistants, signaling real-world fit beyond benchmark optics. Its scale and open availability make it a natural fit for cost-sensitive engineering teams that need very long context, custom fine-tuning, or sovereign on-premise deployment without per-token API constraints.

LLM Gatewaytencent/kimi-k3kimi-k3

Quick Info

Powered by
Provider
LLM Gateway
Model key
tencent/kimi-k3
Release date
Jul 16, 2026
Last updated
Jul 16, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$3.00
Output token cost
$15.00

Limits

Output tokens
1,048,576 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare Kimi K3 (Tencent Cloud) pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Kimi K3 (Tencent Cloud)

LLM Gateway

CoverageBenchmark

A MindStudio blog post dated September 2, 2026 reports on Tencent's internal blind evaluation in which 163 internal experts scored Hy4 preview, GLM 5.3, and Kimi K3 against each other on 203 real engineering tasks. Kimi K3 received an average score of 2.94 and lost to Hy4 preview in 51.2% of head-to-head comparisons (w The benchmark data positions Kimi K3 as a top-tier open-weight baseline against which new Chinese frontier models are measured. While Kimi K3 lost more head-to-heads than it won against Hy4 preview, its average score of 2.94 is competitive and close to Hy4's 2.99 and GLM 5.3's 2.92. The narrow margins across all three

LLM Gateway

Coverage

Moonshot AI paused new subscriptions to Kimi K3 on July 19, 2026, just 48 hours after launch, after user demand pushed the startup's GPU clusters to capacity; existing subscribers were unaffected and the company said it is adding compute. The pause coincided with plans to release the model's open weights on July 27, sh Kimi K3 is a 2.8-trillion-parameter mixture-of-experts model, the largest open-weight AI system released to date by total parameter count, containing 896 specialized expert subnetworks with 16 activated per token under a component called Stable LatentMoE. That sparse-routing architecture keeps per-token compute compara

LLM Gateway

CoverageBenchmark

InferenceX profiles Kimi K3 as Moonshot AI's flagship open-weights model, the first open 3T-class model, with 2.8T total parameters, 104B active, and a 1M-token context window. The architecture uses one dense layer followed by a hybrid stack: 68 Kimi Delta Attention layers and 24 Gated MLA layers, with a Top-16/898 MoE routing pattern mixing KDA and gated MLA at a 3:1 ratio. Moonshot's official X post and tech blog confirm K3 was live on Kimi.com, Kimi Work, Kimi Code, and the Kimi API from July 16, 2026, with open weights promised by July 27. The model card lists native multimodality for text, image, and video, plus MXFP4 quantization support, and ships under a custom Moonshot Kimi K3 License rather than MIT or Apache. Positioning is explicitly agentic and coding-first, reflecting Moonshot's framing of roughly 2.5× scaling efficiency versus Kimi K2.5. Stable LatentMoE routing effectively activates 16 of 896 experts per token, enabling the 2.8T-parameter footprint while keeping active compute at 104B.

LLM Gateway

Coverage

Moonshot AI marked Kimi K3 Open Day by releasing the K3 model weights, the K3 technical report, and supporting infrastructure technologies MoonEP, FlashKDA, and AgentEnv. The excerpt explicitly describes K3 as a 2.8-trillion-parameter Mixture-of-Experts model with native visual understanding and a 1-million-token context window, with roughly three times the parameters of K2.5. Weights are available for download for research, development, and end-user product integration under the Kimi K3 License. The technical report details a 3:1 mix of Kimi Delta Attention and Gated MLA for efficient long-context modeling, with block-level attention residuals improving cross-layer information flow. Stable LatentMoE activates 16 of 896 routed experts, using SiTU-GLU and Quantile Balancing for training stability under extreme sparsity. Moonshot reports a 2.5× improvement in scaling efficiency under constrained compute via innovations including KDA, Attention Residuals, and MoonEP, with each unit of compute producing roughly 2.5× as much intelligence as before.

LLM Gateway

Coverage

Layer3 Labs' buyer's guide confirms Moonshot AI released the full Kimi K3 weights on July 27, 2026, on schedule, as a true open-weights release downloadable from Hugging Face and GitHub rather than an API-only offering. Kimi K3 is identified as the successor to Kimi K2 with 2.8 trillion parameters, billed by Moonshot a The model ships under the Kimi K3 License rather than MIT, carrying two notable conditions: any Model-as-a-Service business with aggregate revenue over $20 million in any 12-month period must sign a separate agreement with Moonshot before commercial use, and products with more than 100 million monthly active users or m

LLM Gateway

Coverage

A "China Pulse" newsletter article dated August 20, 2026 reports that Moonshot AI released Kimi K3 in late July 2026 as a 2.8-trillion-parameter open-weight model distributed free of charge, making it the largest model any lab has open-sourced to date — surpassing Meta's Llama and earlier Chinese releases. The model su The same article notes that the Kimi K3 release preceded Alibaba's Qwen3.8 by approximately one week, setting off a parameter race among Chinese labs in the open-weight tier. In a follow-up update, the piece reports Moonshot AI's decision to take a 30% revenue cut from enterprise customers accessing Kimi K3 weights — t

LLM Gateway

Coverage

TrendForce reports that Moonshot AI opened API access for its flagship Kimi K3 on July 16, 2026 and formally released the full open-source weights on July 27, 2026, positioning it as the largest open-source model in the world at that time. The release was characterized as a market-shaking event comparable to DeepSeek R Kimi K3 is described as a 2.8-trillion-parameter MoE model that activates about 104.2 billion parameters per token, with 896 experts, 16 activated per token (sparsity of roughly 56) and 2 shared experts. The piece places Kimi K3 in context against subsequent open-source releases including DeepSeek V4 Flash (284B total

Videos about Kimi K3 (Tencent Cloud)

More models around Kimi K3 (Tencent Cloud)