MiMo-V2-Flash is a Mixture-of-Experts language model engineered to balance high-level intelligence with operational efficiency. It utilizes a sparse architecture featuring 309 billion total parameters, with only 15 billion active parameters per inference, allowing it to maintain speed while handling demanding tasks. The model is built on a novel hybrid attention architecture that interleaves a 128-token sliding window with full attention in a 5:1 ratio. This design choice significantly reduces KV-cache storage requirements, making the model particularly effective for long-context applications and complex reasoning scenarios where performance and resource management are critical.
The model incorporates Multi-Token Prediction to enhance its generative capabilities, positioning it as a strong competitor in software engineering and scientific reasoning benchmarks. Its design lineage focuses on practical utility, showing top-tier performance in coding evaluations like SWE-bench and scientific assessments such as GPQA-Diamond. By optimizing the interaction between its 256 experts, the model achieves a lightweight footprint relative to its total parameter count, offering a versatile tool for developers and researchers who require a robust foundation for agentic workflows and general-purpose assistance.