Sulat.com
AI models
InferX logo

Model details

mimo-v25

MiMo-V2.5-Pro is a 1.02-trillion-parameter Mixture-of-Experts model from Xiaomi's AI division, activating roughly 42 billion parameters per token through a stack of 70 layers (one dense layer plus 69 MoE layers) and a hidden size of 6,144. Released in late April 2026 as the first open-weight entry in the MiMo Pro tier, it ships with publicly downloadable weights under the MIT license, removing commercial restrictions and allowing fine-tuning, redistribution, or paid hosting without caveats. Xiaomi's MiMo team, led by ex-DeepMind-affiliated researcher Luo Fuli, framed the architecture around efficiency for long-context workloads rather than raw scale alone.

Two design choices define the model's practical character. A hybrid attention scheme blends Sliding Window Attention and Global Attention at roughly a 6-to-1 ratio with a 128-token window, reportedly yielding about a sevenfold reduction in KV-cache size. Three lightweight Multi-Token Prediction modules sit alongside the main stack, providing approximately three times faster autoregressive output. The native FP8 (E4M3) mixed-precision format further lowers serving cost, making the model a fit for teams that need a very large open-weight model with predictable long-context behavior, while the open-weights posture invites community quantisation and reproduction work that was not possible with earlier API-only MiMo Pro variants.

InferXmimo-v25mimo

Quick Info

Powered by
Provider
InferX
Model key
mimo-v25
Release date
Apr 22, 2026
Last updated
Apr 22, 2026
Knowledge cutoff
2024-12
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
100,000 tokens
Context window
1,000,000 tokens

Latest news about mimo-v25

Videos about mimo-v25

More models around mimo-v25