MiMo V2.5 Pro is built around a mixture-of-experts design with roughly one trillion total parameters and about forty-two billion activated per token, paired with a hybrid attention layout that interleaves sliding window and global attention at a six-to-one ratio. This combination lets the model handle very long contexts efficiently, reducing key-value cache storage by nearly seven times while preserving coherence across extended inputs. A multi-token prediction module with three lightweight prediction heads accelerates inference and is designed to speed up reinforcement learning rollouts, making the model practical for production-scale serving and continued training alike.
Positioned for demanding agentic and software engineering workloads, MiMo V2.5 Pro sustains complex multi-step trajectories spanning thousands of tool calls while maintaining instruction following over its million-token context window. Xiaomi reports benchmark gains over its predecessor, including a Coding Agent score of fifty-seven point two, SWE-Bench Pro at seventy-three point seven, and MiMo Coding Bench at sixty-eight point four, reflecting strong performance on realistic engineering tasks. Open weights give teams the ability to self-host, fine-tune, and integrate the model into custom agentic pipelines where long-horizon reasoning, tool use, and reliable instruction following are essential.