Currently listed through these providers:
Model details
MiMo M2.5
MiMo M2.5 is a Mixture-of-Experts language model at the center of Xiaomi's v2.5 family, designed to handle long-context reasoning and tool-augmented workflows without the per-token compute cost of a comparably sized dense model. Its hybrid attention layer weaves together sliding-window and full-attention mechanisms to keep KV-cache memory manageable at the million-token scale, making extended document reviews, codebase analyses, and multi-step agent trajectories genuinely practical. A multi-token prediction head boosts output tokens per inference step, compressing latency in streaming scenarios. As a native omnimodal model, it ingests text, images, audio, and video while producing text output, with vision and file input built into the same interface developers interact with for tool calling and reasoning.
Positioned as the standard tier within the family, MiMo M2.5 prioritizes long-context efficiency and balanced per-token cost over the maximum depth that its Pro sibling delivers at higher compute expense. It carries Pro-level agentic performance while keeping inference costs roughly half of what the deeper variant demands, making it well-suited for developer workflows that move between document understanding and coding tasks. The standard tier excels where implicit prompt caching and consistent per-call efficiency matter more than pushing every query to its theoretical reasoning ceiling. Vision, file input, and tool calling integrate natively, targeting workflows that require reliable multimodal comprehension without the overhead of always invoking the broader parameter pool.
Quick Info
Powered by- Provider
- Vercel AI Gateway
- Model key
- xiaomi/mimo-v2.5
- Release date
- Apr 22, 2026
- Last updated
- Apr 22, 2026
- Knowledge cutoff
- 2024-12
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.14
- Output token cost
- $0.28
Limits
- Output tokens
- 131,100 tokens
- Context window
- 1,050,000 tokens
Transparent token rates
Compare MiMo M2.5 pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about MiMo M2.5
No articles yet. Fetch the latest news to show it here.
