OpenRouter
Xiaomi has officially started the transition from its MiMo-V2 series to the newer MiMo-V2.5 model family, marking an important update for developers using
Model details
MiMo-V2.5 is designed as a native full-modal model that ingests images, video, audio, and text together within a one-million-token context window, allowing cross-modal reasoning over long, mixed-format inputs rather than treating each modality as a separate pipeline. Xiaomi framed the release as a step up from the MiMo-V2 series, positioning V2.5 at what the company calls the Pareto frontier of capability and token efficiency, meaning stronger answers at lower cost per useful output. On community and router-level leaderboards, MiMo-V2.5 has drawn substantial real-world traffic, and a parallel Xiaomi announcement described a trillion-parameter MiMo-V2.5-Pro sibling with about 42B active parameters and agent performance benchmarked against leading frontier systems, signaling the family’s push into autonomous task execution.
In practical use, MiMo-V2.5 fits teams that need an open-weights foundation model for agentic workflows, multimodal document and media understanding, and long-context retrieval or analysis without paying frontier-tier rates. The model supports tool calling, structured outputs, and tunable sampling, so it slots into existing orchestration layers for browsing, reasoning, and operating across external services, and its trillion-parameter-capable family design suggests headroom for heavier deployments when more capacity is required. For builders who want a balance of multimodal perception, reasoning depth, and operating efficiency on self-hostable weights, MiMo-V2.5 is positioned as a flexible backbone that can scale from everyday chat and coding helpers to more demanding agent pipelines.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
OpenRouter
Xiaomi has officially started the transition from its MiMo-V2 series to the newer MiMo-V2.5 model family, marking an important update for developers using
OpenRouter
A community-authored NVIDIA developer-forums thread (a3refaat, May 2026) documents a reproducible recipe for serving a third-party NVFP4 quantization of MiMo-V2.5 (lukealonso/MiMo-V2.5-NVFP4) on a 2× DGX Spark / GB10 cluster. The setup uses vLLM 0.21.1rc1.dev39 on CUDA 13.2 with tensor parallelism = 2, instanttensor lo Three concrete correctness fixes are detailed: (1) the ModelOpt mixed-precision path must dispatch MXFP8 linear layers to the MXFP8 method rather than falling through when the checkpoint mixes MXFP8 dense layers with NVFP4 experts; (2) the weight_scale_inv field is UE8M0 MXFP8 scale metadata and should not be reciproca
OpenRouter
XIAOMI-W (01810.HK) announced that the Xiaomi MiMo-V2.5 series models officially launched public beta testing, alongside optimization of Token Plan pr..., Provide HK Stocks News and Financial News, including Stocks’ company news, result, world economic data, world markets news, china’s policy, warrant and CBBC news
OpenRouter
OpenRouter's LLM rankings (data through Aug 19, 2026) show Xiaomi's MiMo-V2.5 holding the #3 spot on the all-models leaderboard with 6.16T tokens processed, behind DeepSeek V4 Flash 0731 (11.2T) and Tencent Hy3 (9.47T). The 33% week-over-week change indicator is the strongest momentum signal in the set, suggesting rapi The rankings page also breaks down task-level leadership, and MiMo-V2.5's aggregate token throughput places it ahead of GPT-5.6 Luna (5.73T), GLM 5.2 (4.37T), and Claude Opus 5 (2.58T). For developers evaluating routing and fallback strategies on OpenRouter, this is a current, first-party data point that Xiaomi's model
OpenRouter
Compare MiMo-V2.5 from Xiaomi to other AI models on key metrics including benchmarks, price, context length, and other model features.
OpenRouter
MiMo-V2.5 is a native omnimodal model by Xiaomi. $0.14 per million input tokens, $0.28 per million output tokens. 1,048,576 token context window, maximum output of 131,072 tokens. Higher uptime with 2 providers. Includes independent benchmarks from Artificial Analysis.
This exact model name is also listed by 17 other providers.