Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Ambient logo

Model details

MiMo-V2.5

Xiaomi's MiMo-V2.5 family is positioned around agentic, long-horizon work, with the flagship Pro variant described as capable of independently completing professional tasks that would take human experts days or weeks and of sustaining more than a thousand tool calls in a single run. That emphasis on autonomy is paired with a roughly one-million-token context window, which the OpenRouter listing reports as 1M tokens and which is intended to let the model ingest large codebases, document collections, or extended agent traces without losing coherence. Benchmark-wise, Xiaomi highlights top rankings on ClawEval, GDPVal, and SWE-bench Pro, framing the model as a strong generalist for complex software engineering alongside broader agentic evaluation, and the family also includes a lighter Flash tier that community users have begun running on multi-node NVIDIA GB10 setups.

Because the weights for MiMo-V2.5-Pro are published on Hugging Face under the XiaomiMiMo organisation, the line is genuinely open-weight rather than API-only, which makes it attractive for teams that want to self-host, fine-tune, or wire the model into custom agent frameworks. In practice, that combination of open weights, a very large context window, and explicit tuning for tool calling and software engineering tasks points to a model best suited to research assistants, autonomous coding agents, and retrieval-heavy pipelines where the model has to plan across many steps and external services. Buyers comparing against purely proprietary frontier models should weigh the Pro variant's reported agentic benchmark results and long-context behaviour against the availability of weights they can deploy on their own infrastructure.

Ambientxiaomi/mimo-v2.5mimo

Quick Info

Powered by
Provider
Ambient
Model key
xiaomi/mimo-v2.5
Release date
Apr 22, 2026
Last updated
Apr 22, 2026
Knowledge cutoff
2024-12
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.40
Output token cost
$2.00

Limits

Output tokens
131,072 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare MiMo-V2.5 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about MiMo-V2.5

OpenRouter

CoverageBenchmark

A community-authored NVIDIA developer-forums thread (a3refaat, May 2026) documents a reproducible recipe for serving a third-party NVFP4 quantization of MiMo-V2.5 (lukealonso/MiMo-V2.5-NVFP4) on a 2× DGX Spark / GB10 cluster. The setup uses vLLM 0.21.1rc1.dev39 on CUDA 13.2 with tensor parallelism = 2, instanttensor lo Three concrete correctness fixes are detailed: (1) the ModelOpt mixed-precision path must dispatch MXFP8 linear layers to the MXFP8 method rather than falling through when the checkpoint mixes MXFP8 dense layers with NVFP4 experts; (2) the weight_scale_inv field is UE8M0 MXFP8 scale metadata and should not be reciproca

Videos about MiMo-V2.5

More models around MiMo-V2.5