MiMo-V2-Flash is a Mixture-of-Experts foundation model built by Xiaomi with 309 billion total parameters and 15 billion active parameters per forward pass. Its architecture uses a hybrid attention mechanism with 64 attention heads and 8 key-value heads, layered across 48 transformer blocks with RMS normalization and SwigLU activation. The model incorporates 256 expert neurons in its MoE layers, activating 8 per token, which allows it to dynamically route computations for efficiency. Position encoding uses Rotary Position Embedding with a theta of 640,000, and the tokenizer operates with a vocabulary of 151,680 tokens. The architecture also supports a hybrid-thinking toggle that lets users control whether the model engages in explicit step-by-step reasoning or produces direct responses.
The model achieved global top-1 ranking among open-source models on both SWE-bench Verified and SWE-bench Multilingual benchmarks, delivering performance on par with Claude Sonnet 4.5 at roughly 3.5% of the cost. Its strengths are particularly evident in software engineering tasks, mathematical reasoning, and agent scenarios where tool use and extended context matter. Released under an MIT license in December 2025, MiMo-V2-Flash is available as open weights, allowing developers to run and fine-tune it independently. Compared to the proprietary MiMo v2 Pro variant, the Flash version offers a larger context window, built-in reasoning mode, and native tool-calling capabilities, making it well-suited for developers seeking an open, cost-effective foundation for coding assistants, autonomous agents, and complex problem-solving applications.