Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Umans AI Coding Plan logo

Model details

Kimi K3

Kimi K3 is a 2.8-trillion-parameter open-weight model that combines Kimi Delta Attention with an Attention Residuals design, scaling MoE sparsity through a Stable LatentMoE framework that activates 16 out of 896 experts. It introduces native vision capability alongside text understanding and extends to a one-million-token context window, positioning itself as the first openly available 3T-class model. The architecture aims at frontier intelligence for long-horizon coding, structured knowledge work, and multi-step reasoning, rather than narrow chat tasks, and it is shipped with open weights for self-hosting and audit.

According to Moonshot AI's own evaluations, Kimi K3 reaches frontier-level results on their suite while still trailing the strongest proprietary systems, and it consistently outperforms other models they tested. In practice it is well suited to agentic coding pipelines, where its long context, tool calling, structured output support, and reasoning capacity let a single model drive multi-step development workflows. Hosting providers like Umans expose it through API endpoints that integrate with popular agent frameworks such as Claude Code, Cursor, OpenCode, and Zed, making the open weights usable inside familiar developer tools without giving up on per-token control.

Umans AI Coding Planumans-kimi-k3kimi-k3

Quick Info

Powered by
Provider
Umans AI Coding Plan
Model key
umans-kimi-k3
Release date
Jul 16, 2026
Last updated
Jul 16, 2026
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
131,072 tokens
Context window
1,048,576 tokens

Latest news about Kimi K3

Umans AI

Coverage

We0.ai confirms that on July 27, 2026, Moonshot AI released full Kimi K3 weights on Hugging Face, GitHub, and ModelScope, alongside the technical report, Kimi K3 License, and deployment guidance. The release marks the largest openly available weights in the 3-trillion-parameter class. The model is described as a native multimodal MoE with about 104 billion activated parameters per token and a 1-million-token context window. Disclosed materials cover architecture, post-training, long-context RL infrastructure, expert-parallel training, inference optimizations, benchmarks, and engineering case studies.

Umans AI Coding Plan

CoverageBenchmark

The LLM List aggregator corroborates that Umans AI Coding Plan hosts Kimi K3 as `umans-kimi-k3` at `https://api.code.umans.ai/v1` with a 1.05M-token context window and 131K max output, listed as a free endpoint. Kimi K3 itself is documented with ~2,779.9B parameters, native Kimi K3 architecture, and support for reasoni Benchmark and pricing context shows Umans's free endpoint contrasting with paid providers like CrofAI ($2/$8 per 1M tokens) and NanoGPT ($2/$10), while Artificial Analysis-derived scores place Kimi K3 at Terminal-Bench Hard 82.7 (max effort) and Terminal-Bench 2.1 at 46.0, with ~36–37 tok/s throughput and TTFT around 3

Umans AI

Coverage

Geopolitechs reposted Moonshot AI's 'Kimi K3 Open Day' announcement from 27 July 2026, confirming release of the Kimi K3 model weights, the technical report, and supporting infrastructure technologies MoonEP, FlashKDA, and AgentEnv. Moonshot describes K3 as its most capable model: a 2.8-trillion-parameter Mixture-of-Ex Technical report details include KDA and Gated MLA mixed at a 3:1 ratio for efficient long-context modeling, block-level attention residuals for cross-layer information flow, and Stable LatentMoE with 16/896 routed experts using SiTU-GLU and Quantile Balancing for training stability. The visual encoder MoonViT-V2 is tr

Umans AI

CoverageBenchmark

WhatLLM confirms Kimi K3 launched as a hosted Moonshot AI service on 16 July 2026, with Moonshot subsequently publishing the full checkpoint, a custom Kimi K3 License, deployment recipes, and technical report. The model is a 2.8-trillion-parameter sparse MoE activating 16 of 896 routed experts per token, with native vi Pricing is documented at $3/M uncached input, $0.30/M cached input, and $15/M output—significantly higher than Kimi K2.6. WhatLLM recommends testing K3 for complex coding, large repositories, document-heavy research, and visual production while keeping cheaper models for routine traffic. The article flags that most det

Umans AI

Coverage

Nathan Lambert's Interconnects analysis describes Moonshot AI releasing Kimi K3 on 16 July 2026 as a 2.8T-parameter MoE model with full weights promised for 27 July 2026. Lambert argues this narrows the open-to-closed and US-to-China model performance gap from a debated 6–9 months to roughly 3–5 months, and frames K3 a Benchmark rankings cited include #2 overall on Vals AI, #3 on Artificial Analysis's Intelligence Index (behind only Claude Fable and GPT-5.6 Sol Max while being cheaper), and #1 overall in Frontend Code Arena. Lambert characterizes Moonshot as going toe-to-toe with Anthropic and OpenAI with far fewer resources, and cal

Umans AI

Coverage

Moonshot AI released Kimi K3 on 16 July 2026 as a 2.8-trillion-parameter mixture-of-experts model with native vision, a one-million-token context window, and only 16 of 896 experts active per token, per the Beam AI model directory review. Architectural features include Kimi Delta Attention and Attention Residuals to ma Benchmarks show K3 scoring 57 on Artificial Analysis's Intelligence Index, reaching a GDPval-AA v2 Elo of 1668 (above GPT-5.5 and Claude Opus 4.8 in that test), leading AutomationBench-AA at 53%, and placing second on AA-Briefcase, indicating strong agentic task completion. Beam estimates roughly $0.94 per completed be

Umans AI

Coverage

Gizmochina reports Moonshot AI launched Kimi K3 as the largest open-weights model to date with 2.8 trillion parameters, claiming frontier-level capability. The model is available now via the Kimi iOS and Android apps, kimi.com, and the Kimi Work desktop client on a free tier with usage limits. Kimi K3 uses a Mixture-of-Experts architecture with 16 of 896 experts activated per request, a 1 million-token context window, native image support (and reportedly video), plus Kimi Delta Attention for efficient long-context handling. API pricing starts at $3 per million input tokens according to the article.

Umans AI

CoverageBenchmark

BenchLM's aggregator page lists Kimi K3 as released on 16 July 2026 with a 1.05M-token context window, reasoning capability, and weight access marked pending. The model scores 74.6/100 in the composite, ranking 7 of 232 tracked models, with its strongest eligible category being Multimodal & Grounded at rank 1. Capabili API pricing is listed at $3 per million input tokens, $15 per million output tokens, and $0.30 per million cached input tokens, with a blended cost around $9. Reported throughput is 35 tokens/second with a first-token latency of 61.49 seconds. The dashboard notes 48 published benchmark rows with some tracked slots empt

Umans AI

CoverageBenchmark

CodingFleet compiles a head-to-head comparison of Kimi K3 against Anthropic's Claude Opus 5 using BenchLM and Artificial Analysis data. Kimi K3 details confirmed include MoE architecture with 2.8T total parameters, 16 of 896 experts active (roughly 32B active), and a 1,048,576-token context window. Kimi K3's listed API pricing is $3 per million input tokens, $15 per million output, and $0.30 per million cached input, with open-weights licensing under a Modified MIT. The comparison gives Opus 5 a higher aggregate score (85.88 vs 79.98) while positioning Kimi K3 as a cheaper, competitive open-weights alternative.

Videos about Kimi K3

More models around Kimi K3