Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Pioneer logo

Model details

Kimi K3

Kimi K3 is a 2.8-trillion-parameter open model from Moonshot AI, described by its creators as the first open model to reach the 3T class. It is designed around Kimi Delta Attention and Attention Residuals, paired with a Stable LatentMoE framework that activates 16 of 896 experts, yielding roughly 2.5 times better scaling efficiency than the previous Kimi K2 generation. The architecture also brings native vision capabilities alongside text, letting the model accept visual input without a separate pipeline, and the wider Kimi family has held the upper bound of open-model parameter sizes for nine of the past twelve months leading up to the release.

In practice, Kimi K3 targets frontier-level work on long-horizon coding, broad knowledge tasks, and reasoning. Moonshot's evaluation suite places it behind the top proprietary systems Claude Fable 5 and GPT 5.6 Sol while still outperforming every other model tested, signaling strong but not top-tier open performance. The model is positioned to sustain extended engineering sessions with minimal oversight, making it a fit for developers and teams who want an open, large-scale model for agentic coding, multi-step reasoning, and multimodal understanding without depending on a closed provider.

Pioneermoonshotai/Kimi-K3-Fastkimi-k3

Quick Info

Powered by
Provider
Pioneer
Model key
moonshotai/Kimi-K3-Fast
Release date
Jul 16, 2026
Last updated
Jul 16, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$4.50
Output token cost
$22.50

Limits

Output tokens
131,072 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare Kimi K3 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Kimi K3

Pioneer

Coverage

The Kimi K3 technical report (alphaXiv, August 2026) introduces a 2.8T-parameter Mixture-of-Experts model with 104B activated parameters, native vision, and a 1-million-token context window. Architecture combines Kimi Delta Attention and Attention Residuals for better information flow across sequence length and depth, Post-training combines reinforcement learning across general, agentic, and coding domains at multiple reasoning-effort levels, targeting compositional generalization and long-horizon execution. Evaluations show frontier-level performance on long-horizon coding, agentic, knowledge, reasoning, and vision tasks; overall t

Pioneer

Coverage

On July 27, 2026 — Moonshot AI's declared Kimi K3 Open Day — the lab released the Kimi K3 model weights alongside its technical report and the supporting training/inference infrastructure (MoonEP, FlashKDA, AgentEnv). Kimi K3 is described as a 2.8-trillion-parameter Mixture-of-Experts model with native visual understan This candidate is the closest available to a first-party source among the supplied excerpts, reposting Moonshot's own Open Day announcement translated from a WeChat post, and provides the most authoritative architectural detail on Kimi K3 — directly supporting the model/family focus criterion. It confirms that the open

Pioneer

CoverageAnalysis

Tahir's Medium architecture breakdown focuses on the efficiency story beneath Moonshot AI's Kimi K3, which carries 896 experts and activates only 16 per token — roughly 1.8% of the model at any time. Compared against other open MoE models, this is the smallest activation ratio in the comparison set: Nemotron 3 Ultra ac The piece frames Kimi K3's significance not as benchmark dominance but as a genuine shift in inference economics and semiconductor demand implications from extreme sparsity at frontier scale. It positions K3 alongside DeepSeek's earlier catch-up moment and argues that Chinese models are reshaping what is possible at th

Pioneer

CoverageBenchmark

Yotta Labs provides a hardware-oriented technical breakdown of Moonshot AI's Kimi K3, confirming 2.8 trillion total parameters in a sparse MoE (896 experts with 16 active per token, ~104B active parameters), a 1-million-token context window, and natively integrated text + vision modalities with always-on reasoning. The The article also places Kimi K3 in ecosystem context as Moonshot's successor to the K2 line, arriving days before Alibaba's Qwen 3.8-Max preview, and notes architectural innovations including Kimi Delta Attention aimed at making the 1M context window affordable to serve, plus Moonshot's claimed 2.5× scaling-efficiency

Pioneer

CoverageBenchmark

WhatLLM's write-up profiles Kimi K3 as Moonshot AI's flagship multimodal reasoning model launched 16 July 2026 as a hosted service: 2.8 trillion total parameters in a sparse MoE activating 16 of 896 routed experts per token, with a one-million-token context window and native vision. It is positioned for long-horizon ag The pricing listed is $3/M uncached input, $0.30/M cached input, and $15/M output, well above Kimi K2.6. WhatLLM ranks K3 first on Arena's blind frontend coding leaderboard at launch and competitive across Moonshot's coding and agentic suite, while Moonshot's own launch post concedes K3 still trails Claude Fable 5 and

Pioneer

Coverage

Moonshot AI released Kimi K3 on 16 July 2026 as a 2.8-trillion-parameter mixture-of-experts model with native vision and a one-million-token context window; only 16 of 896 routed experts activate per token, with Kimi Delta Attention and Attention Residuals intended to make that scale practical for long-horizon tasks. F On independent benchmarks Artificial Analysis gave K3 a 57 on its Intelligence Index, a GDPval-AA v2 Elo of 1668 (above GPT-5.5 and Claude Opus 4.8, behind Claude Fable 5), and led AutomationBench-AA at 53% while placing second on AA-Briefcase. Estimated cost was about $0.94 per completed benchmark task. The review pos

Pioneer

Coverage

Nathan Lambert's ecosystem analysis on Interconnects argues that Kimi K3, released 16 July 2026 as a 2.8T-parameter MoE with weights scheduled for 27 July, is the closest an open model has come to the frontier since DeepSeek R1. The core thesis is that the open-to-closed or US-to-China model performance gap has likely Lambert catalogs Kimi K3's launch-time rankings: 2 overall on the Vals AI index, 3 on Artificial Analysis's Intelligence Index (behind Claude Fable and GPT-5.6 Sol Max while being cheaper), 1 on Frontend Code Arena, and other strong agentic results. He frames Moonshot as executing on standard scaling levers (data, algo

Pioneer

CoverageBenchmark

LLM-Stats tracks Kimi K3 on a composite LLM Stats Score of 53.0 at a blended price of $3.39 per million tokens, placing it between DeepSeek-V4.1-Flash (51.8) and GPT-6 Astra (59.6). The scorecard reports an AA-Briefcase Elo of 1548 (#2) from Artificial Analysis and a MathVision score of 0.98 (#1) sourced from the Kimi Per-conversation-depth tracking on LLM-Stats shows K3 averaging 15.8 on turn 1 (95% CI 15.9–16.4, 6 of 114), dropping to 14.7 across turns 2–10 (31 of 115, 95% CI 14.9–15.4), and 14.6 across turns 11–30 (11 of 100, 95% CI 14.8–15.7), a cumulative decay of −1.2σ with no data past turn 30. Subcategory signals include Gam

Videos about Kimi K3

More models around Kimi K3