MiMo-V2.6-Pro-UltraSpeed is a variant within Xiaomi's MiMo-V2.6 series, rolled out alongside the open-sourced MiMo-V2.6-Pro and MiMo-V2.6-Flash models. Where Pro and Flash emphasize capability and efficiency respectively, the UltraSpeed build targets users who need extreme generation speed, with Xiaomi describing it as delivering up to 20x faster output at the same quality as the base Pro model. The broader series is natively omnimodal, and Xiaomi framed the work as part of its exploration of the RSI path, scaling reinforcement learning compute on verifiable, complex tasks so that the model can expand its capability frontier through exploration and feedback rather than purely supervised training.
The practical appeal of the UltraSpeed variant is straightforward: teams that already want MiMo-V2.6-Pro's quality but are bottlenecked on token throughput can opt into a faster inference profile without giving up response fidelity. Xiaomi's training approach for the series emphasized large-batch asynchronous RL, with Flash and Pro each completing 30 RL steps across roughly 750,000 trajectories in under six days during a publicly observed production run, and Xiaomi reported that reinforcement learning continued improving throughout the run and generalized beyond the training distribution. That combination of speed-optimized serving and a training pipeline grounded in verifiable-task RL makes UltraSpeed a fit for latency-sensitive agentic or interactive workloads where output volume is the limiting factor.