MiMo-V2.5-Pro-UltraSpeed is a speed-optimized variant within Xiaomi's open-weights family, positioned as a serving configuration for high-throughput agent workloads rather than a standalone model. Its underlying engine is the MiMo-V2.5-Pro-FP4-DFlash backbone, which applies MXFP4 quantization specifically to the MoE experts while keeping the rest of the network at higher precision. This expert-focused quantization shrinks model size and memory-bandwidth pressure with what the technical card describes as near-lossless quality, making it practical to deploy trillion-parameter decoding on commodity hardware.
To push throughput further, the system pairs that quantized backbone with a BF16 DFlash drafter that uses block-diffusion speculative decoding, proposing an entire block of tokens per forward pass so the main model can verify them in a single step. Together, these two techniques attack both dominant costs of large-scale inference: per-parameter bit width and the number of backbone forward passes. The deployment reportedly runs on a single 8-GPU commodity node using Xiaomi's purpose-built TileRT inference engine, with no custom silicon required. This combination enables claimed output speeds above 1,000 tokens per second, which Xiaomi frames as roughly 5–15 times faster than contemporary frontier models in independent reporting. Access is currently structured as an extended closed beta through the Xiaomi MiMo API Platform, with priority given to professional teams building coding agents and similar latency-sensitive workflows.