Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
OpenRouter logo

Model details

MiMo-V2.5

MiMo-V2.5 is designed as a native full-modal model that ingests images, video, audio, and text together within a one-million-token context window, allowing cross-modal reasoning over long, mixed-format inputs rather than treating each modality as a separate pipeline. Xiaomi framed the release as a step up from the MiMo-V2 series, positioning V2.5 at what the company calls the Pareto frontier of capability and token efficiency, meaning stronger answers at lower cost per useful output. On community and router-level leaderboards, MiMo-V2.5 has drawn substantial real-world traffic, and a parallel Xiaomi announcement described a trillion-parameter MiMo-V2.5-Pro sibling with about 42B active parameters and agent performance benchmarked against leading frontier systems, signaling the family’s push into autonomous task execution.

In practical use, MiMo-V2.5 fits teams that need an open-weights foundation model for agentic workflows, multimodal document and media understanding, and long-context retrieval or analysis without paying frontier-tier rates. The model supports tool calling, structured outputs, and tunable sampling, so it slots into existing orchestration layers for browsing, reasoning, and operating across external services, and its trillion-parameter-capable family design suggests headroom for heavier deployments when more capacity is required. For builders who want a balance of multimodal perception, reasoning depth, and operating efficiency on self-hostable weights, MiMo-V2.5 is positioned as a flexible backbone that can scale from everyday chat and coding helpers to more demanding agent pipelines.

OpenRouterxiaomi/mimo-v2.5mimo

Quick Info

Powered by
Provider
OpenRouter
Model key
xiaomi/mimo-v2.5
Release date
Apr 22, 2026
Last updated
Apr 22, 2026
Knowledge cutoff
2024-12
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.14
Output token cost
$0.28

Limits

Output tokens
131,072 tokens
Context window
1,050,000 tokens

Transparent token rates

Compare MiMo-V2.5 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about MiMo-V2.5

OpenRouter

Coverage

Xiaomi has officially started the transition from its MiMo-V2 series to the newer MiMo-V2.5 model family, marking an important update for developers using

OpenRouter

CoverageBenchmark

A community-authored NVIDIA developer-forums thread (a3refaat, May 2026) documents a reproducible recipe for serving a third-party NVFP4 quantization of MiMo-V2.5 (lukealonso/MiMo-V2.5-NVFP4) on a 2× DGX Spark / GB10 cluster. The setup uses vLLM 0.21.1rc1.dev39 on CUDA 13.2 with tensor parallelism = 2, instanttensor lo Three concrete correctness fixes are detailed: (1) the ModelOpt mixed-precision path must dispatch MXFP8 linear layers to the MXFP8 method rather than falling through when the checkpoint mixes MXFP8 dense layers with NVFP4 experts; (2) the weight_scale_inv field is UE8M0 MXFP8 scale metadata and should not be reciproca

OpenRouter

CoveragePreview

XIAOMI-W (01810.HK) announced that the Xiaomi MiMo-V2.5 series models officially launched public beta testing, alongside optimization of Token Plan pr..., Provide HK Stocks News and Financial News, including Stocks’ company news, result, world economic data, world markets news, china’s policy, warrant and CBBC news

OpenRouter

Official sourceBenchmark

OpenRouter's LLM rankings (data through Aug 19, 2026) show Xiaomi's MiMo-V2.5 holding the #3 spot on the all-models leaderboard with 6.16T tokens processed, behind DeepSeek V4 Flash 0731 (11.2T) and Tencent Hy3 (9.47T). The 33% week-over-week change indicator is the strongest momentum signal in the set, suggesting rapi The rankings page also breaks down task-level leadership, and MiMo-V2.5's aggregate token throughput places it ahead of GPT-5.6 Luna (5.73T), GLM 5.2 (4.37T), and Claude Opus 5 (2.58T). For developers evaluating routing and fallback strategies on OpenRouter, this is a current, first-party data point that Xiaomi's model

OpenRouter

Official sourceComparison

Compare MiMo-V2.5 from Xiaomi to other AI models on key metrics including benchmarks, price, context length, and other model features.

OpenRouter

Official sourceBenchmark

MiMo-V2.5 is a native omnimodal model by Xiaomi. $0.14 per million input tokens, $0.28 per million output tokens. 1,048,576 token context window, maximum output of 131,072 tokens. Higher uptime with 2 providers. Includes independent benchmarks from Artificial Analysis.

Videos about MiMo-V2.5

More models around MiMo-V2.5