Currently listed through these providers:
Model details
MiMo-V2.5
MiMo-V2.5 has emerged as a fresh entry in the broader MiMo lineup, drawing immediate attention from enthusiasts running compact NVIDIA DGX Spark and GB10 workstations. Community members in the dedicated DGX Spark user forum flagged it as a new model shortly after launch and noted that some early benchmarks looked promising enough to fit within a two-node Spark setup, hinting at a weight footprint aimed at high-end desktop inference rather than hyperscale clusters. The same community space has since grown into an active discussion thread, reflecting real practitioner interest in evaluating the family on consumer-grade accelerated hardware.
A closely related Omni-branded sibling, MiMo V2.5 Omni, has been demonstrated running across three DGX Spark nodes using tensor parallelism together with multi-token prediction, sustaining around 39 tokens per second at a one-the cataloged API limit. That deployment profile, tagged as an agentic-AI project, suggests the V2.5 generation is being explored for long-context assistant and tool-using workloads where extended memory and steady throughput matter more than raw single-request latency. Together, the new-model announcement and the Omni long-context demonstration point to a family oriented toward accessible, locally hosted reasoning rather than purely cloud-scale serving.
Quick Info
Powered by- Provider
- CrossModel
- Model key
- xiaomi/mimo-v2.5
- Release date
- Apr 22, 2026
- Last updated
- Apr 22, 2026
- Knowledge cutoff
- 2024-12
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.16
- Output token cost
- $0.32
Limits
- Output tokens
- 128,000 tokens
- Context window
- 1,000,000 tokens
Latest news about MiMo-V2.5
Videos about MiMo-V2.5
More models around MiMo-V2.5
This exact model name is also listed by 21 other providers.