Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Novita AI logo

Model details

MiMo-V2.6-Pro

MiMo-V2.6-Pro is the flagship release in Xiaomi's MiMo-V2.6 series, described by the company as its most capable model to date and built as a natively omnimodal system that handles text along with image, audio, and video inputs in a single architecture. The series was released and open-sourced on September 22nd, 2026, marking a key step in Xiaomi's exploration of a recursive self-improvement path in which reinforcement learning compute is scaled on verifiable, complex tasks so the model can extend its capability frontier through exploration and feedback. Alongside the Pro model, Xiaomi introduced a Flash variant aimed at balancing intelligence, efficiency, and cost, as well as a Pro-UltraSpeed variant that targets up to 20x faster output at comparable quality for latency-sensitive deployments.

Independent benchmark reporting from Xiaomi positions MiMo-V2.6-Pro as a strong contender in agentic and coding evaluations, with scores of 71.9 on DeepSWE v1.1, 26.5 on ProgramBench, and 63.2 on the in-house MiMo Code Bench, results that place it ahead of the Flash sibling and well above the prior MiMo-V2.5-Pro generation. The model is designed for practical agentic workflows rather than narrow chat use, with reported performance on additional agent benchmarks such as Toolathlon-verified, GDPVal 2.1 AA, Automation Bench v1.0.6, and Agents' Last Exam reinforcing a focus on tool use, automation, and long-horizon task completion. Open-weight availability makes MiMo-V2.6-Pro well suited to teams that want to self-host a frontier-grade omnimodal model for coding assistants, research agents, or complex multi-step automation pipelines.

Novita AIxiaomimimo/mimo-v2.6-promimo

Quick Info

Powered by
Provider
Novita AI
Model key
xiaomimimo/mimo-v2.6-pro
Release date
Sep 22, 2026
Last updated
Sep 22, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.435
Output token cost
$0.87

Limits

Output tokens
131,072 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare MiMo-V2.6-Pro pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about MiMo-V2.6-Pro

Deep Infra

CoverageBenchmark

Xiaomi released MiMo V2.6 on September 22, 2026, a family of open-weight, natively omnimodal language models that handle text, images, video, and audio with a 1 million-token context window. Built around the slogan "Scaling Reinforcement Learning Toward Self-Improvement," the lineup includes Pro, Flash, Pro UltraSpeed, and a 9B distill, all under an MIT license and available on Hugging Face, ModelScope, Xiaomi's API, and OpenRouter. Pro is a 1.02-trillion-parameter sparse MoE model with 42B active parameters, Flash is a 309B/15B-active variant, and UltraSpeed targets low latency. Xiaomi reports Pro leading peers on AutomationBench v1.0.6 (53.1), Terminal Bench 2.1 (89.9), and CyberGym (94.0), while trailing on harder coding and exploit benchmarks like Terminal Bench 4.0 (34.9). Artificial Analysis places Pro at an Intelligence Index of 46 at roughly $0.13 per task, with OpenRouter pricing of $0.14/$0.28 per million tokens for Flash and $0.435/$0.87 for Pro.

Deep Infra

CoverageBenchmark

On 22 September 2026, Xiaomi released its MiMo-V2.6 model family as open-weight checkpoints under an MIT licence, with the flagship MiMo-V2.6-Pro topping Artificial Analysis's open-weight Intelligence Index at a score of 46. The release comes from Xiaomi's MiMo team and bundles four models together with a technical report and open-source RL tooling on GitHub. Xiaomi frames the launch around a single mixed reinforcement-learning run across coding, agents, visual tasks and cybersecurity, replacing earlier domain-specific training. MiMo-V2.6-Pro is a 1.02T-parameter sparse MoE with 42B active parameters, a 1M-token context, and text, image, video and audio input, with a 573 GB checkpoint. API pricing sits at $0.435 per million input tokens and $0.87 per million output tokens, while the Flash sibling (309B total, 15B active) is priced at $0.14 and $0.28, and an UltraSpeed variant costs ten times more. Self-hosting Pro requires datacentre GPUs such as 8x H200, whereas Flash fits on a single multi-GPU node and a 9B Qwen distill runs on a laptop. Xiaomi was named in Anthropic's 10 September 2026 distillation report alongside six other China-based labs, though no link to V2.6 has been documented.

Deep Infra

Coverage

Sebastian Raschka describes MiMo-V2.6-Pro as the current leader among open-weight models on weighted-average benchmarks, attributing gains to data and post-training recipe rather than architectural novelty. Its backbone uses grouped query attention with a small 128-token sliding-window attention window, keeping the design a classic transformer configuration. Key training improvements highlighted in the Xiaomi technical report include expanding agent tasks and training across multiple harnesses, raising DeepSWE pass@1 on held-out harnesses from roughly 50% to 66%, and replacing a simple correctness verifier with an agentic grader that inspects execution traces. Reinforcement learning batches of 1,568 prompts times 16 rollouts yielding 25,088 trajectories and 2.7 to 3.7 billion training tokens per update drove much of the gain.

Deep Infra

CoverageBenchmark

BenchLM's October 8, 2026 snapshot places MiMo-V2.6-Pro at a capability score of 74.1/100, ranking 11th of 216 tracked models with coding at rank 11 and agentic at rank 13. The model is listed at $0.43 input, $0.87 output, $0.004 cached input, and a blended $0.65 per million tokens. Speed is measured at 46 tokens per second with a 48.61-second first-token time and a 1M-token context window, positioning the model near the top of its open-weight cohort on coding and agentic benchmarks. Coding draws on 4 verified benchmarks at 64.6 (rank 11 of 146, 93rd percentile), while the 9-benchmark agentic category scores 66.9 at the 90th percentile, making coding its strongest published category.

Deep Infra

CoverageDocumentation

Xiaomi officially released and open-sourced the MiMo-V2.6 series, framing it as a step along a recursive self-improvement path built on verifiable tasks and scaled reinforcement-learning compute. V2.6-Pro scores 46 on the Artificial Analysis Intelligence Index, surpassing Kimi K3 and Qwen3.8 Max to become the most powerful open-source model, though it still trails leading closed models. Both Pro and Flash are native multimodal models, and V2.6 series pricing matches V2.5 series pricing, Xiaomi states, pushing out the intelligence-versus-cost Pareto frontier. MiMo-V2.5-Pro and MiMo-V2.5 requests will auto-route to V2.6 on October 14, 2026, at 18:00 UTC+8, with old model IDs becoming invalid after October 21, 2026, at 10:00 UTC+8.

Deep Infra

CoverageBenchmark

llm-stats lists MiMo-V2.6-Pro at $0.435 per million input tokens and $0.870 per million output tokens via Xiaomi, with reused prompt prefixes at $0.0036 per million cached input tokens and a 1.0M context window. The model has 1.0T MoE parameters under a commercial-use MIT license, with documentation hosted at platform.xiaomimimo.com and weights on Hugging Face. Operational telemetry from October 2 to October 8, 2026, shows a p95 time-to-first-token of about 5.31 seconds and a sustained output floor of 145 characters per second. Quality tracker scores the model at 14.5 on turn 1 and 13.3 on turns 2–10 across 56 to 106 evaluated models, with capability tiers placing MiMo-V2.6-Pro at the upper end of the open-weight Frontier (200B+) cohort.

Deep Infra

CoverageBenchmark

Artificial Analysis rates MiMo-V2.6-Pro at 46 on its Intelligence Index, placing it well above the open-weight comparable median of 18, and ranks it among the strongest models in intelligence. The model supports text, image, speech, and video input with text output, a 1M-token context window, and 1.0T total / 42B active parameters under an MIT license. MiMo-V2.6-Pro pricing is listed at $0.43 per 1M input tokens and $0.87 per 1M output tokens with a 99% cache discount bringing cached input to about $0.13 per 1M tokens. Output throughput is measured at 42.5 tokens per second, with the evaluator noting it is notably slow and somewhat verbose compared to peers. The metrics land it in the top quartile for intelligence and third quartile for cost within its open-weight size class.

Deep Infra

Coverage

Xiaomi released MiMo-V2.6-Pro and MiMo-V2.6-Flash under the MIT license, posting weights to Hugging Face on September 21, 2026, with the official announcement dated September 22. Pro is a 1.02-trillion-parameter MoE model with 42B active parameters across 70 layers, while Flash has 309B total / 15B active parameters. Both checkpoints natively handle text, image, video, and audio and support a 1-million-token context window. The architecture combines sliding-window and global attention layers with a 681M-parameter vision encoder and two audio encoders. Xiaomi recommends SGLang with speculative decoding and vLLM examples use tensor parallelism of 8 for Pro and 4 for Flash. Xiaomi's benchmark tables place V2.6-Pro at 82.0 on OSWorld-Verified, 76.9 on Toolathlon-Verified, and 53.1 on AutomationBench, approaching Claude Opus 5 and GPT-5.6 Sol on agentic coding and computer-use tasks.

Videos about MiMo-V2.6-Pro

More models around MiMo-V2.6-Pro