Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Xiaomi logo

Model details

MiMo-V2.5-Pro-UltraSpeed

MiMo-V2.5-Pro-UltraSpeed is a speed-optimized variant within Xiaomi's open-weights family, positioned as a serving configuration for high-throughput agent workloads rather than a standalone model. Its underlying engine is the MiMo-V2.5-Pro-FP4-DFlash backbone, which applies MXFP4 quantization specifically to the MoE experts while keeping the rest of the network at higher precision. This expert-focused quantization shrinks model size and memory-bandwidth pressure with what the technical card describes as near-lossless quality, making it practical to deploy trillion-parameter decoding on commodity hardware.

To push throughput further, the system pairs that quantized backbone with a BF16 DFlash drafter that uses block-diffusion speculative decoding, proposing an entire block of tokens per forward pass so the main model can verify them in a single step. Together, these two techniques attack both dominant costs of large-scale inference: per-parameter bit width and the number of backbone forward passes. The deployment reportedly runs on a single 8-GPU commodity node using Xiaomi's purpose-built TileRT inference engine, with no custom silicon required. This combination enables claimed output speeds above 1,000 tokens per second, which Xiaomi frames as roughly 5–15 times faster than contemporary frontier models in independent reporting. Access is currently structured as an extended closed beta through the Xiaomi MiMo API Platform, with priority given to professional teams building coding agents and similar latency-sensitive workflows.

Xiaomimimo-v2.5-pro-ultraspeedmimobeta

Quick Info

Powered by
Provider
Xiaomi
Model key
mimo-v2.5-pro-ultraspeed
Release date
Jun 8, 2026
Last updated
Jun 9, 2026
Knowledge cutoff
2024-12
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$1.305
Output token cost
$2.61

Limits

Output tokens
131,072 tokens
Context window
1,048,576 tokens

Latest news about MiMo-V2.5-Pro-UltraSpeed

Xiaomi

Coverage

On June 24, 2026, 36kr/ZDXX reported that Xiaomi's MiMo Open Platform extended the limited-time chat and API access window for its MiMo-V2.5-Pro-UltraSpeed model, originally launched on June 9, 2026 and scheduled to close on June 23. According to the Xiaomi MiMo team's official notice, the extension was triggered by de The same report details MiMo-V2.5-Pro-UltraSpeed's architecture and performance as cited from Xiaomi: it is a Mixture-of-Experts model jointly developed by the Xiaomi MiMo team and the TileRT AI inference systems team, with 1 trillion total parameters, approximately 42 billion activated parameters per forward pass, and

Xiaomi

Coverage

Third-party reporting from Yahoo/Tech (dated June 8, 2026) provides the most concrete technical breakdown of how MiMo-V2.5-Pro-UltraSpeed achieves its claimed 1,000+ tokens/s. The deployment reportedly runs on a single 8-GPU commodity node with no custom silicon, relying on TileRT as a purpose-built inference engine an The article contextualizes the speed against Artificial Analysis figures: GPT-5.5 at roughly 68 tokens/s, Claude Opus 4.6 around 71, Haiku near 98, and Gemini Flash at about 192 tokens/s, framing UltraSpeed as roughly 5–15x faster depending on baseline. It also notes Cerebras hit 969 tokens/s on Llama 3.1 405B and Groq

Xiaomi

Coverage

Second-source third-party coverage of the June 8–9, 2026 UltraSpeed launch specifies the underlying MiMo-V2.5-Pro as a 1.02T-parameter MoE model with 42B active parameters and native 1-million-token context support. It uses hybrid attention with a sliding-window to global attention ratio of 6:1, plus three layers of Mu Pricing rose from ¥6 to ¥18 per million output tokens (3x) in exchange for an order-of-magnitude throughput boost. Compared to mainstream closed-source flagships—Claude Opus 4.6 around 40 t/s and GPT-5.4 high-speed at 120–180 t/s—Xiaomi claims the first trillion-scale model past 1,000 t/s on commodity GPUs, while custo

Videos about MiMo-V2.5-Pro-UltraSpeed

More models around MiMo-V2.5-Pro-UltraSpeed