Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
FastRouter logo

Model details

Kimi K2

Kimi K2 is built on a mixture-of-experts architecture, meaning it routes problems through specialized subnetworks rather than activating the entire model at once. This design lets it punch far above its weight—32 billion active parameters backed by a trillion total parameters—while keeping inference costs low. The model is explicitly engineered for agentic work: it doesn't just answer queries, it takes action, invoking tools dynamically and autonomously executing complex, multi-step tasks. Open weights mean anyone can download and adapt it, and its benchmark results across software engineering, multilingual coding, and mathematics show it competing with models that cost far more to operate.

The model ships in two forms: a foundation version for researchers who want full control for fine-tuning, and an instruction-tuned variant ready for drop-in use in chat and agentic applications. The instruction variant is reflex-grade, optimized for speed without requiring long thinking processes. Its tool-use capabilities let developers give it a set of functions and describe a goal; Kimi K2 automatically figures out how to chain calls together to get there. Compared to models like Claude Opus 4 that cost $15 per million input tokens, Kimi K2's pricing makes it attractive for production agentic pipelines at scale.

FastRoutermoonshotai/kimi-k2kimi-k2

Quick Info

Powered by
Provider
FastRouter
Model key
moonshotai/kimi-k2
Release date
Jul 11, 2025
Last updated
Jul 11, 2025
Knowledge cutoff
2024-10
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.55
Output token cost
$2.20

Limits

Output tokens
32,768 tokens
Context window
131,072 tokens

Transparent token rates

Compare Kimi K2 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Kimi K2

DevPass (LLM Gateway)

CoverageBenchmark

A July 13, 2026 third-party technical explainer on Lorphic provides a detailed architectural breakdown of the base Kimi K2 model, attributing its specifications to Moonshot's official technical documentation and Hugging Face model cards. K2 is described as a 1-trillion-parameter Mixture-of-Experts (MoE) model with 32 b The same Lorphic explainer clarifies that the Kimi K2 family (K2, K2.5, K2.6, K2.7 Code) shares this architectural foundation but each variant carries meaningfully different capabilities, licensing terms, and recommended use cases. The post walks through version-by-version differences, API setup considerations, and ben

Videos about Kimi K2

More models around Kimi K2