Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Melious logo

Model details

Kimi K3

Kimi K3 is an open-weight, natively multimodal reasoning model created by Moonshot AI, with its weights published under the moonshotai organization on Hugging Face alongside a dedicated license file and an accompanying K3 tech report PDF. Moonshot positions it as their most capable model to date and the first open 3T-class release, scaled to 2.8T parameters within a Stable LatentMoE framework that activates 16 out of 896 experts. The architecture introduces Kimi Delta Attention paired with Attention Residuals, a design aimed at efficient long-context reasoning while keeping inference compute manageable for such a large sparse model.

Practically, Kimi K3 targets frontier workloads in long-horizon coding, knowledge work, and agentic tool use, where its the cataloged API limit context, native vision input, and tool-calling support make it well suited to navigating large repositories, iterating against logs, tests, and runtime feedback, and debugging visually grounded tasks. The combination of strong reasoning capability, structured output, and open availability makes it a flexible foundation for teams building coding assistants, research agents, and complex multi-step workflows that need both deep context retention and the ability to call external tools.

Meliouskimi-k3kimi-k3

Quick Info

Powered by
Provider
Melious
Model key
kimi-k3
Release date
Jul 16, 2026
Last updated
Jul 16, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$3.1878
Output token cost
$15.939

Limits

Output tokens
131,072 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare Kimi K3 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Kimi K3

Melious

CoverageBenchmark

Moonshot AI released Kimi K3 on July 16, 2026, as a 2.8 trillion parameter sparse Mixture-of-Experts model with 896 experts (16 active per token, ~104B active parameters), a 1 million token context window, and native text-and-vision integration with always-on reasoning. According to Yotta Labs, the open weights shipped The piece highlights that K3 does not fit on a single GPU or node — running it requires a multi-node cluster with 1.6 TB+ of aggregate GPU memory — and notes Moonshot's claimed 2.5x scaling-efficiency improvement over K2 via Kimi Delta Attention. Third-party benchmarks cited include strong coding leaderboard performanc

Melious

CoverageBenchmark

WhatLLM corroborates that Kimi K3 launched as a hosted service on July 16, 2026, with 2.8 trillion total parameters — the first announced open model in the three-trillion-parameter class. As a sparse MoE transformer, only 16 of 896 routed experts activate per token. The model targets agent-style workloads: navigating l Moonshot's own launch post acknowledges K3 still trails Claude Fable 5 and GPT-5.6 Sol overall, but reaches the same frontier band on many coding and agentic tasks. API pricing is $3/M uncached input, $0.30/M cached input, and $15/M output — far more than K2.6. WhatLLM recommends K3 for complex coding, huge repositorie

Melious

Coverage

Nathan Lambert's analysis positions Kimi K3's July 16, 2026 release by Moonshot AI as a true frontier model and the strongest open model ever released — the closest open models have been to the frontier since DeepSeek R1. K3 is described as a 2.8T parameter MoE model with weights promised for July 27. Lambert argues th The article cites independent benchmark placements: #2 on Vals AI index, #3 on Artificial Analysis's Intelligence Index (beaten only by Claude Fable and GPT-5.6 Sol Max while being cheaper), and #1 on Frontend Code Arena. Lambert frames this as a Chinese lab going toe-to-toe with Anthropic and OpenAI with far fewer res

Melious

Coverage

BenchLM confirms the July 27, 2026 open-weight release: 96 Safetensors shards, the Kimi K3 License, and the technical report are now live on the Hugging Face model repository. The model is officially documented as 2.8T total / 104B active parameters, with exactly 1,048,576 tokens of context, accepting text and images. Moonshot's launch post recommends 64 or more accelerators for serving due to communication and expert routing overhead at this scale, making the direct API the practical testing path for most teams. API pricing tracks at $3.00/M cache-miss input, $0.30/M cache-hit input, and $15.00/M output, flat across the full contex

Melious

CoverageBenchmark

Simon Willison confirms Moonshot AI announced Kimi K3 on July 16, 2026, describing it as their "most capable model to date, with 2.8 trillion parameters." The model became available via Moonshot's website and API immediately, with an open weight release promised by July 27, 2026. Moonshot positioned K3 as the first "op K3's pricing of $3/million input and $15/million output tokens is noted as the most expensive model released by a Chinese AI lab, matching Anthropic's Claude Sonnet series and a significant increase over K2.6's $0.95/$4. Willison's firsthand qualitative test using OpenRouter to generate an SVG pelican cost 25 cents (95

Videos about Kimi K3

More models around Kimi K3