Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Qiniu logo

Model details

Xiaomi/Mimo-V2-Flash

MiMo-V2-Flash is a Mixture-of-Experts foundation model built by Xiaomi with 309 billion total parameters and 15 billion active parameters per forward pass. Its architecture uses a hybrid attention mechanism with 64 attention heads and 8 key-value heads, layered across 48 transformer blocks with RMS normalization and SwigLU activation. The model incorporates 256 expert neurons in its MoE layers, activating 8 per token, which allows it to dynamically route computations for efficiency. Position encoding uses Rotary Position Embedding with a theta of 640,000, and the tokenizer operates with a vocabulary of 151,680 tokens. The architecture also supports a hybrid-thinking toggle that lets users control whether the model engages in explicit step-by-step reasoning or produces direct responses.

The model achieved global top-1 ranking among open-source models on both SWE-bench Verified and SWE-bench Multilingual benchmarks, delivering performance on par with Claude Sonnet 4.5 at roughly 3.5% of the cost. Its strengths are particularly evident in software engineering tasks, mathematical reasoning, and agent scenarios where tool use and extended context matter. Released under an MIT license in December 2025, MiMo-V2-Flash is available as open weights, allowing developers to run and fine-tune it independently. Compared to the proprietary MiMo v2 Pro variant, the Flash version offers a larger context window, built-in reasoning mode, and native tool-calling capabilities, making it well-suited for developers seeking an open, cost-effective foundation for coding assistants, autonomous agents, and complex problem-solving applications.

Qiniuxiaomi/mimo-v2-flashmimo

Quick Info

Powered by
Provider
Qiniu
Model key
xiaomi/mimo-v2-flash
Release date
Dec 16, 2025
Last updated
Feb 4, 2026
Knowledge cutoff
2024-12-01
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.10
Output token cost
$0.30

Limits

Output tokens
256,000 tokens
Context window
256,000 tokens

Transparent token rates

Compare Xiaomi/Mimo-V2-Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Xiaomi/Mimo-V2-Flash

Qiniu

Coverage

The Baidu encyclopedia entry on Xiaomi MiMo documents MiMo-V2-Flash as a 309 billion total parameter / 15 billion active parameter mixture-of-experts model released alongside the MiMo-Embodied model in December 2025 under the MIT license, with base weights published on Hugging Face. The model was specifically designed MiMo-V2-Flash was released as part of Xiaomi's broader MiMo model lineage, which originated in April 2025 as a 7B reasoning model and grew under the leadership of Luo Fuli, who joined Xiaomi from DeepSeek in November 2025 to head the AI Large Model Team. Subsequent MiMo milestones referenced include HySparse hybrid spa

Videos about Xiaomi/Mimo-V2-Flash

More models around Xiaomi/Mimo-V2-Flash