MiMo-V2.5-Pro is a sparse Mixture-of-Experts model with over a trillion total parameters and 42 billion active parameters during inference. It inherits a hybrid-attention architecture and three-layer Multi-Token Prediction from the MiMo-V2-Flash backbone, enabling it to handle long-horizon reasoning and complex software engineering challenges. This design makes it particularly strong at agentic workflows and tasks that demand sustained coherence over extended contexts. According to Xiaomi's published benchmarks, MiMo-V2.5-Pro tied for first place among open-weights models on the Artificial Analysis Intelligence Index, demonstrating that efficient expert routing can deliver frontier-level performance without activating the full parameter count on every token.
The model was released alongside the standard MiMo-V2.5 version, both open-sourced under the MIT license for commercial use. XiaomI released both variants with weights available on Hugging Face, enabling enterprises and independent developers to run them locally or in private cloud environments. MiMo-V2.5-Pro excels at agentic tasks such as powering claw-style task systems and benchmark environments like SWE-bench Pro, where it scored competitively against leading proprietary models. With a one-million-token context window and significantly lower operational costs than comparable open-weights alternatives, it targets teams that need high benchmark performance on coding and agentic work without frontier-model pricing.