Xiaomi Token Plan (Singapore)
One-time purchase unlocks both MiMo-V2-Pro and Omni flagship models, plus TTS model free across all tiers for a limited time. Unleash powerful productivity with Xiaomi MiMo
Model details
MiMo-V2.5-Pro is positioned as Xiaomi's flagship entry in the MiMo family, built around a trillion-scale Mixture-of-Experts architecture with roughly 1T total parameters and approximately 42B active per token, paired with a context window that stretches to 1M tokens. That combination of sparse activation and very long context is what lets the model sustain complex software engineering sessions and long-horizon agent workflows without losing track of earlier code, tool calls, or system instructions. Xiaomi highlights strong general agentic capability, and external benchmark listings point to top placements on agent- and engineering-oriented evaluations including ClawEval, GDPVal, and SWE-bench Pro, with the model described as able to independently complete professional-grade tasks involving more than a thousand tool calls.
In practical terms, MiMo-V2.5-Pro is a fit for teams building autonomous coding agents, multi-step research assistants, or any pipeline that chains dozens to hundreds of tool invocations while keeping a large project state in view. Its open weights are published on Hugging Face under the XiaomiMiMo organization, which means it can be self-hosted, fine-tuned, or audited rather than treated as a closed API endpoint, and community discussion already shows it running on multi-node consumer accelerators. The broader V2.5 series adds sharper instruction following and improved omni-modal understanding alongside more natural speech synthesis, and Xiaomi has been actively migrating users forward by deprecating the prior V2 series in favor of V2.5, while bundling promotional access to Pro and Omni flagships with free TTS for limited windows.
A provider subscription or plan supersedes token-based pricing for this model.
Xiaomi Token Plan (Singapore)
One-time purchase unlocks both MiMo-V2-Pro and Omni flagship models, plus TTS model free across all tiers for a limited time. Unleash powerful productivity with Xiaomi MiMo
Xiaomi Token Plan (Singapore)
A developer on the NVIDIA DGX Spark / GB10 forum published a reproducible recipe for running the community-quantized lukealonso/MiMo-V2.5-NVFP4 checkpoint on a 2× DGX Spark cluster with TP=2, using vLLM 0.21.1rc1.dev39 on a CUDA 13.2 build. The configuration uses instanttensor load format, Triton attention with DIFFKV, The post documents three concrete correctness fixes that were required to get the mixed-precision checkpoint running: dispatching MXFP8 linear layers to the MXFP8 method rather than letting them fall through, treating weight_scale_inv as UE8M0 MXFP8 scale metadata (not reciprocal-inverted), and using a QKV-aware tensor
This exact model name is also listed by 17 other providers.