Kimi K3 is a 2.8 trillion-parameter Mixture-of-Experts language model developed by Moonshot AI, with parameters activated sparsely so the model can deliver a large overall capacity without paying the full inference cost on every token. The architecture is built on Kimi Delta Attention and Attention Residuals, paired with a Stable LatentMoE framework that activates 16 of 896 experts per pass, which Moonshot credits with roughly a 2.5x improvement in scaling efficiency compared with its predecessor. These structural changes, combined with refined training and data recipes, let the lab convert compute into capability more effectively, an important consideration for a team working under tighter hardware constraints than many Western competitors.
The model was announced on July 16, 2026 and its open-weight files followed on July 27, 2026, making it the strongest openly available release at the time. Independent evaluations place it near the top of language and coding leaderboards, including a second-overall position on the Vals AI index, third on Artificial Analysis (only behind Claude Fable 5 and GPT 5.6 Sol, and at a lower price), and first on Frontend Code Arena. Practically, Kimi K3 is aimed at developers and organizations that want frontier-class reasoning, long-context handling, and customization freedom without being locked into a single provider, while still offering a hosted API for teams that prefer managed serving.