Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Ollama Cloud logo

Model details

kimi-k3

Kimi K3 is distributed through Ollama's cloud catalog as a vision-capable model, tagged alongside tools and thinking features, which positions it for multimodal workflows that mix image inputs with reasoning and structured tool calls. A long-form analysis published by Zvi Mowshowitz on both LessWrong and Substack frames the model in terms of strong benchmark performance and its place among recent open-weight releases, describing it as a very capable system while situating it several months behind the closed model frontier. Independent commentary notes the model's raw capability stands out among open options, with the gap between pre-training and closed-source frontier models characterized as wider than the gap in post-training refinement.

As hosted on Ollama Cloud, Kimi K3 fits into a platform that emphasizes open model weights, open-source code, and user data privacy, with cloud regions available in the United States, Europe, and Singapore. The combination of vision input, tool calling, and reasoning-oriented thinking modes makes it suitable for agent-style pipelines, code-assistant integrations, and analytical tasks that benefit from structured outputs. Practitioners seeking a transparent open-weight alternative for multimodal reasoning tasks will find Kimi K3 a practical choice, particularly when pairing its capabilities with Ollama's ecosystem of CLI and editor integrations.

Ollama Cloudkimi-k3kimi-k3

Quick Info

Powered by
Provider
Ollama Cloud
Model key
kimi-k3
Release date
Jul 16, 2026
Last updated
Jul 27, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$3.00
Output token cost
$15.00

Limits

Output tokens
131,072 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare kimi-k3 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about kimi-k3

Ollama Cloud

Coverage

The alphaXiv-hosted abstract of the Kimi K3 technical report describes the model as a 2.8T-parameter Mixture-of-Experts with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. The architecture centers on Kimi Delta Attention and Attention Residuals for improved informati Extensive evaluations show Kimi K3 achieving frontier-level performance on long-horizon coding, agentic, knowledge, reasoning, and vision tasks, while its overall performance still trails the most powerful proprietary models, namely Claude Fable 5 and GPT-5.6 Sol. The report explicitly states Moonshot releases the full

Ollama Cloud

Coverage

Geopolitechs reproduces Moonshot AI's Kimi K3 Open Day announcement from July 27, 2026, confirming the release of K3 model weights, the technical report, and three key supporting infrastructure technologies: MoonEP, FlashKDA, and AgentEnv. Kimi K3 is described as a 2.8-trillion-parameter Mixture-of-Experts model with n The technical report covers KDA + AttnRes (3:1 ratio of KDA to Gated MLA for efficient long-context, with block-level attention residuals improving cross-layer information flow), Stable LatentMoE (16 of 896 routed experts activated per token, with SiTU-GLU and Quantile Balancing for stability), and MoonViT-V2 (a visual

Ollama Cloud

CoverageAnalysis

Medium architecture analysis from July 24, 2026, highlights Kimi K3's 896 routed experts with only 16 activated per token — an activation ratio of about 1.8% of the model at any time, the smallest among major open models (compared to Nemotron 3 Ultra at 4.3%, Mixtral 8x7B at 3.1%, and Deepseek V3 at 3.1%). The piece ar The analysis frames K3 as a model whose significance is efficiency rather than raw scale, suggesting the architecture reshapes inference costs and semiconductor demand for Chinese AI deployments. The piece positions K3 as evidence that Chinese labs are not just catching up to OpenAI and Anthropic on benchmarks but are

Ollama Cloud

Coverage

Moonshot AI released Kimi K3 on July 16, 2026, as a 2.8-trillion-parameter Mixture-of-Experts model, with full weights scheduled for public release on July 27, 2026, according to Nathan Lambert's analysis on Interconnects. The piece frames K3 as the closest an open-weights model has come to the frontier since DeepSeek The analysis situates K3 within a broader ecosystem reflection, suggesting that if a Chinese lab can approach closed-model performance and release its weights, scarcity and pricing power across the sector are less secure. Lambert notes that any contribution from adversarial distillation from US frontier models would be

Ollama Cloud

CoverageBenchmark

Whatllm.org's July 20, 2026 article documents that Kimi K3 launched through Kimi products and APIs on July 16, with Moonshot subsequently publishing the full checkpoint, a custom Kimi K3 License, deployment recipes, and a technical report. Pricing is confirmed at $3/M uncached input, $0.30/M cached input, and $15/M out K3 is described as best suited for complex coding, large repositories, document-heavy research, and visual production, with notable competitive results including 1st place in Arena's blind frontend coding ranking at launch. A one-million-token context, native vision, and long-horizon tool use position K3 as an agent-cl

Ollama Cloud

Coverage

Beam AI's model directory entry describes Moonshot AI releasing Kimi K3 on July 16, 2026, as a 2.8-trillion-parameter mixture-of-Experts model with native vision and a one-million-token context window, activating 16 of 896 experts per token. The review details that on July 17, Z.ai fell roughly 28% and a competing Chin Benchmark coverage places K3 at 57 on Artificial Analysis's Intelligence Index, with a GDPval-AA v2 Elo of 1668 (above GPT-5.5 and Claude Opus 4.8, behind Claude Fable 5), a leading AutomationBench-AA score of 53%, and second place on AA-Briefcase. Artificial Analysis estimated roughly $0.94 per completed benchmark tas

Ollama Cloud

CoverageBenchmark

Kylon's benchmark breakdown covers Kimi K3's July 16, 2026 release as a 2.8 trillion-parameter, open-weight model with a 1 million-token context window, native vision, and an always-on reasoning mode. Within 24 hours of launch, K3 climbed to #1 on Arena's frontend coding leaderboard and #3 on Artificial Analysis's Inte On the Artificial Analysis Intelligence Index across 189 models, K3 scores 57.1 at the 97th percentile, trailing Claude Fable 5's 60 and GPT-5.6 Sol's 59 but ahead of Claude Opus 4.8 at 57 and GPT-5.5 at 55. The article compares K3 head-to-head with Claude Fable 5, OpenAI's GPT-5.6 Sol, and Zhipu AI's GLM-5.2 across co

Ollama Cloud

CoverageBenchmark

NxCode's coding-agent evaluation guide frames Kimi K3 as having enough benchmark evidence to justify a serious engineering evaluation, but not enough same-harness, independently reproduced evidence to justify a universal best coding model claim. Moonshot reports strong coding results: 67.5 on DeepSWE, 77.8 raw pass rat Artificial Analysis independently places K3 near the frontier but also measures high output-token use, slower-than-median generation, and premium pricing, so capability and operating efficiency point in different directions. The harness is part of the product: K3 is sensitive to preserved thinking history, and Moonshot

Videos about kimi-k3

More models around kimi-k3