Sulat.com
AI models
Hugging Face logo

Model details

MiMo-V2-Flash

MiMo-V2-Flash is a Mixture-of-Experts language model developed by Xiaomi, combining 309 billion total parameters with a sparse activation design that puts roughly 15 billion parameters to work during any given forward pass. This architectural choice allows the model to deliver substantial capability while keeping inference costs manageable. The design centers on a hybrid attention approach that interleaves sliding window attention with global attention in a 5-to-1 ratio, using a 128-token sliding window to achieve near-six-fold reduction in KV-cache storage compared to naive global attention. Multi-Token Prediction extends the model's ability to generate multiple tokens per step, accelerating throughput for high-speed reasoning and agentic workflows. The model also offers a hybrid-thinking toggle, letting users control reasoning behavior through a simple boolean flag.

MiMo-V2-Flash represents Xiaomi's open-source foundation model lineage, released under an open-weight framework that has attracted community tooling including vLLM integration recipes and pipeline configurations. On software engineering benchmarks—specifically SWE-bench Verified and SWE-bench Multilingual—the model ranks as the top-performing open-source option globally, delivering performance that rivals Claude Sonnet 4.5 while operating at roughly 3.5% of the cost. A 256K context window supports extended reasoning chains and multi-step agent tasks. The design philosophy prioritizes practical agent scenarios, coding tasks, and complex reasoning chains, positioning the model as a versatile backbone for both research and production applications requiring open-weight access and cost-efficient deployment.

Hugging FaceXiaomiMiMo/MiMo-V2-Flashmimo

Quick Info

Powered by
Provider
Hugging Face
Model key
XiaomiMiMo/MiMo-V2-Flash
Release date
Dec 16, 2025
Last updated
Dec 16, 2025
Knowledge cutoff
2024-12
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.10
Output token cost
$0.30

Limits

Output tokens
4,096 tokens
Context window
262,144 tokens

Transparent token rates

Compare MiMo-V2-Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about MiMo-V2-Flash

Hugging Face

Coverage

On December 17, 2025, at Xiaomi's Human-Car-Home Ecosystem Partner Conference, Luo Fuli — newly appointed head of Xiaomi's MiMo large model team — officially unveiled MiMo-V2-Flash, described as the company's next step toward Artificial General Intelligence. The model uses a Mixture-of-Experts architecture and a Hybrid MiMo-V2-Flash also integrates Multi-Token Prediction (MTP), which Xiaomi credits with major efficiency gains during reinforcement-learning training: three-layer MTP reportedly achieves an acceptance length above 3 and roughly 2.5× faster speed on programming tasks, helping address GPU idle time in small-batch on-policy

Videos about MiMo-V2-Flash

More models around MiMo-V2-Flash