Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Novita AI logo

Model details

MiMo-V2.6-Flash

We couldn't load the overview just now. Please try again in a little while.

Novita AIxiaomimimo/mimo-v2.6-flashmimo

Quick Info

Powered by
Provider
Novita AI
Model key
xiaomimimo/mimo-v2.6-flash
Release date
Sep 22, 2026
Last updated
Sep 22, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.14
Output token cost
$0.28

Limits

Output tokens
131,072 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare MiMo-V2.6-Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about MiMo-V2.6-Flash

Deep Infra

CoverageBenchmark

The eesel explainer confirms MiMo V2.6 is a family of open-weight, natively omnimodal models from Xiaomi's MiMo team released on September 22, 2026, with every checkpoint shipping a 1 million-token context window. Flash is explicitly named as a distinct checkpoint in the family, sharing the omnimodal input profile (tex The page frames the release under the tagline "Scaling Reinforcement Learning Toward Self-Improvement," with Xiaomi positioning the series as a step on a recursive self-improvement path via scaling RL compute on verifiable, complex tasks. The author notes that the technical report has enough concrete method behind it t

Deep Infra

Coverage

Xiaomi released and open-sourced the MiMo-V2.6 family on September 21, 2026, and MiMo-V2.6-Flash ships as one of the named variants alongside Pro and a 9B distillation. Per the page, Flash is 309B total parameters with 15B active, omnimodal across text, image, video, and audio, runs a 1M-token context, and was produced On pricing, Flash is listed at $0.14 per million input tokens and $0.28 per million output tokens on Xiaomi's own endpoint, unchanged from V2.5, while a faster Pro UltraSpeed serving variant costs $4.35/$8.70. Xiaomi also shipped the technical report, more than 7,000 RL environments, and the training framework with the

Deep Infra

CoverageBenchmark

Benchable's page for Xiaomi's MiMo-V2.6-Flash, released September 21, 2026, describes a Mixture-of-Experts foundation model with 309B parameters and a 1M-token context, pricing at $0.14 input and $0.28 output per 1M tokens with $0.0028 cache reads. It reports a 99% reliability rate and competitive response times in the 53rd percentile across benchmarks. Highlighted benchmark accuracies include 100% on Hallucinations baseline, 99.5% General Knowledge, 99.0% Email Classification (90th percentile), 92.0% Reasoning, 91.0% Coding, 92.0% Mathematics, and 97.0% Ethics (29th percentile), with 69.0% Instruction Following. The model supports tools, structured outputs, reasoning, response format controls, and accepts text, image, audio, and video input while emitting text.

Deep Infra

CoverageBenchmark

BenchLM's October 8, 2026 snapshot places MiMo-V2.6-Flash at capability 65.9/100, ranking 17th for Coding and 37th for Agentic across 122 to 146 models. Pricing is listed at $0.14 per 1M input and $0.28 per 1M output tokens with $0.003 cache reads and a 1M-token context window. Of 14 published benchmark rows, 9 are verified for Agentic and 3 for Coding; Reasoning, Multimodal, Knowledge, Multilingual, Instruction Following, and Math remain unmeasured. The page reports independent runtime speed as not measured, limiting speed-based conclusions while still marking the model as well-suited for software development and code generation.

Deep Infra

CoverageBenchmark

MiMo-V2.6-Flash ranks 29th globally on RankLLMs with a composite score of 52.4, including a 4th-place OpenCode ranking and 67.2% on SWE-bench Verified. The model runs at 185 tokens per second with sub-220ms time-to-first-token and is reported as 4th in developer adoption on OpenCode with 9,486 daily active developers. Reported benchmark scores include 54.8% on GPQA Diamond, 62.0% on MATH-500, 32.5% on OSWorld, 52.0% on BrowseComp, 64.5% on Terminal-Bench 2.1, and a GDPVal-AA Arena Elo of 1560. The model is an open-weights, open-source release with a 1M-token context window, positioned as an ultra-fast multimodal option.

Deep Infra

Coverage

Xiaomi released MiMo-V2.6-Flash weights on Hugging Face on September 21, 2026, under an MIT license alongside MiMo-V2.6-Pro, with FP8 weights totaling 172.9 GB. The Flash checkpoint has 309B total and 15B active parameters across 48 layers (39 sliding-window and 9 global-attention) with a hidden size of 4096, and includes a 5-layer multi-token-prediction drafter that proposes up to 7 tokens per speculative decoding pass. Both MiMo-V2.6 models are natively multimodal via a 681M-parameter vision encoder and two audio encoders feeding the shared MoE backbone, supporting a 1M-token context. Xiaomi also published a 9B Qwen3.5-9B distillation, a technical report, RL code, and more than 7,000 training environments, and recommends SGLang with speculative decoding or vLLM with tensor parallelism of 4 for Flash.

Deep Infra

CoverageBenchmark

MiMo-V2.6-Flash is an open-weights Xiaomi model scoring 38 on the Artificial Analysis Intelligence Index, placing it well above the 18 median among comparable models. It is priced at $0.14 per 1M input and $0.28 per 1M output tokens with a 98% cache discount, and it runs at 58 tokens per second while remaining reasonably priced relative to peers. The model uses a Mixture-of-Experts design with 309B total and 15B active parameters, ships under an MIT license, and supports a 1M-token context window with text and image input. A reasoning variant is noted, though the page evaluates the reasoning version; Artificial Analysis flags the model as slower than average and very verbose at 240M output tokens on the Intelligence Index.

Videos about MiMo-V2.6-Flash

More models around MiMo-V2.6-Flash