Sulat.com
AI models
ClinePass logo

Model details

Kimi K3

Kimi K3 is presented as Kimi's most capable flagship to date and the first open-source model to reach 2.8 trillion parameters, anchoring Kimi's push toward frontier-scale open weights. It is built on Kimi Delta Attention (KDA), a hybrid linear attention mechanism, paired with Attention Residuals, and ships with native visual understanding alongside a 1M-token context window. Kimi positions this scale and architecture combination for frontier intelligence workloads such as long-horizon coding, broad knowledge work, and advanced problem solving.

Within Kimi's coding lineup, K3 is offered alongside K2.7 Code and exposed through multiple model identifiers that can be selected directly inside Kimi Code clients or third-party coding tools. A recommended 256k variant is highlighted for everyday Q&A, code completion, and routine feature development, while the full 1M variant is suited to longer sessions, with the documentation noting that context compaction may be needed when switching between the two. The combination of an open-weight release, a long context window, and a hybrid attention design makes K3 a practical fit for teams that want a self-runnable frontier-tier model for sustained agentic coding and reasoning tasks.

ClinePasscline-pass/kimi-k3kimi-k3

Quick Info

Powered by
Provider
ClinePass
Model key
cline-pass/kimi-k3
Release date
Jul 16, 2026
Last updated
Jul 16, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$3.00
Output token cost
$15.00

Limits

Output tokens
131,072 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare Kimi K3 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Kimi K3

ClinePass

CoverageAnalysis

Tahir's Medium architecture breakdown (24 July 2026) describes Kimi K3 as a 2.8-trillion-parameter mixture-of-experts model with 896 total experts, of which only 16 are active per token — an activation ratio of about 1.8%. That ratio is lower than comparable sparse models: Nemotron 3 Ultra activates roughly 4.3%, Mixtr The author emphasizes that with 896 experts spread across an estimated 64–72 GPUs, each token must travel between accelerators to reach its assigned experts, making inter-GPU communication the dominant cost rather than raw FLOPs. The piece situates K3 within a broader narrative of Chinese AI labs reshaping inference ec

ClinePass

Coverage

Nathan Lambert analyzes Moonshot AI's 16 July 2026 release of Kimi K3, a 2.8-trillion-parameter mixture-of-experts model whose weights are scheduled for release on 27 July 2026. He frames K3 as a true frontier model and the strongest open-weight release since DeepSeek R1, placing it at #2 overall on the Vals AI index, The post is structured as an ecosystem reflection conditioned on Moonshot honoring the 27 July weight-release promise, contrasting K3's open-weights trajectory with a counterfactual in which China keeps similarly powerful models closed. Lambert notes K3 differs from DeepSeek R1 in being a scaling-execution story rather

ClinePass

CoverageBenchmark

Simon Willison reports that Moonshot AI announced Kimi K3 on 16 July 2026 as their most capable model to date, a 2.8-trillion-parameter system that Moonshot describes as the first "open 3T-class model," with an open-weight release promised by 27 July 2026. Available immediately through Moonshot's website and API, K3 is Willison adds hands-on testing via OpenRouter, noting that K3 accepts image input and generates substantial reasoning output. A pelican-on-a-bicycle SVG prompt consumed 95 input tokens and 16,658 output tokens (13,241 of them reasoning tokens) for a total cost of about 25 cents, and a follow-up image-input run returned

Videos about Kimi K3

More models around Kimi K3