Currently listed through these providers:
Model details
mimo-v25
MiMo-V2.5-Pro is a 1.02-trillion-parameter Mixture-of-Experts model from Xiaomi's AI division, activating roughly 42 billion parameters per token through a stack of 70 layers (one dense layer plus 69 MoE layers) and a hidden size of 6,144. Released in late April 2026 as the first open-weight entry in the MiMo Pro tier, it ships with publicly downloadable weights under the MIT license, removing commercial restrictions and allowing fine-tuning, redistribution, or paid hosting without caveats. Xiaomi's MiMo team, led by ex-DeepMind-affiliated researcher Luo Fuli, framed the architecture around efficiency for long-context workloads rather than raw scale alone.
Two design choices define the model's practical character. A hybrid attention scheme blends Sliding Window Attention and Global Attention at roughly a 6-to-1 ratio with a 128-token window, reportedly yielding about a sevenfold reduction in KV-cache size. Three lightweight Multi-Token Prediction modules sit alongside the main stack, providing approximately three times faster autoregressive output. The native FP8 (E4M3) mixed-precision format further lowers serving cost, making the model a fit for teams that need a very large open-weight model with predictable long-context behavior, while the open-weights posture invites community quantisation and reproduction work that was not possible with earlier API-only MiMo Pro variants.
Quick Info
Powered by- Provider
- InferX
- Model key
- mimo-v25
- Release date
- Apr 22, 2026
- Last updated
- Apr 22, 2026
- Knowledge cutoff
- 2024-12
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 100,000 tokens
- Context window
- 1,000,000 tokens