Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Vercel AI Gateway logo

Model details

MiMo M2.5

MiMo M2.5 is a Mixture-of-Experts language model at the center of Xiaomi's v2.5 family, designed to handle long-context reasoning and tool-augmented workflows without the per-token compute cost of a comparably sized dense model. Its hybrid attention layer weaves together sliding-window and full-attention mechanisms to keep KV-cache memory manageable at the million-token scale, making extended document reviews, codebase analyses, and multi-step agent trajectories genuinely practical. A multi-token prediction head boosts output tokens per inference step, compressing latency in streaming scenarios. As a native omnimodal model, it ingests text, images, audio, and video while producing text output, with vision and file input built into the same interface developers interact with for tool calling and reasoning.

Positioned as the standard tier within the family, MiMo M2.5 prioritizes long-context efficiency and balanced per-token cost over the maximum depth that its Pro sibling delivers at higher compute expense. It carries Pro-level agentic performance while keeping inference costs roughly half of what the deeper variant demands, making it well-suited for developer workflows that move between document understanding and coding tasks. The standard tier excels where implicit prompt caching and consistent per-call efficiency matter more than pushing every query to its theoretical reasoning ceiling. Vision, file input, and tool calling integrate natively, targeting workflows that require reliable multimodal comprehension without the overhead of always invoking the broader parameter pool.

Vercel AI Gatewayxiaomi/mimo-v2.5mimo

Quick Info

Powered by
Provider
Vercel AI Gateway
Model key
xiaomi/mimo-v2.5
Release date
Apr 22, 2026
Last updated
Apr 22, 2026
Knowledge cutoff
2024-12
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.14
Output token cost
$0.28

Limits

Output tokens
131,100 tokens
Context window
1,050,000 tokens

Transparent token rates

Compare MiMo M2.5 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about MiMo M2.5

No articles yet. Fetch the latest news to show it here.

Videos about MiMo M2.5

More models around MiMo M2.5