Moonshot AI (China)
When I saw the update announcement for Kimi K2–0905, my first thought was: finally, a model dares to push context length to 256K. This number is more exaggerated than some startup’s funding pitch …
Model details
Kimi K2 0905 is an open-weight MoE language model that refines its predecessor with a focus on agentic workflows. Built on a trillion-parameter Mixture-of-Experts architecture that activates 32 billion parameters per forward pass, it balances large-scale capability with computational efficiency. The model extends its context handling from the earlier 128k window and is purpose-engineered for developers tackling complex, multi-step tasks requiring tool use, code synthesis, and scaffold-aware reasoning. Its frontend development outputs show measurable improvement in aesthetics and functionality for web, 3D, and related tasks.
This iteration traces its lineage from Kimi K2 0711, carrying forward a training stack that incorporates the MuonClip optimizer for stable large-scale MoE optimization. Benchmark evidence points to competitive performance across LiveCodeBench and SWE-bench for coding tasks, ZebraLogic and GPQA for reasoning, and Tau2 and AceBench for tool-use evaluation. Beyond code, the model maintains strong creative writing capabilities with reduced hallucination compared to earlier versions. Practical design choices like seamless Claude Code compatibility and smooth function calling aim to minimize friction for developers integrating the model into agentic pipelines and development workflows.
Moonshot AI (China)
When I saw the update announcement for Kimi K2–0905, my first thought was: finally, a model dares to push context length to 256K. This number is more exaggerated than some startup’s funding pitch …
Moonshot AI (China)
Kimi K2 0905 is the September update of [Kimi K2 0711](moonshotai/kimi-k2). $0.40 per million input tokens, $2 per million output tokens. 262,144 token context window, maximum output of 262,144 tokens. Higher uptime with 4 providers. Includes independent benchmarks from Artificial Analysis.