Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
EmpirioLabs AI logo

Model details

MiMo V2.5 Pro

MiMo V2.5 Pro sits at the top of Xiaomi's model lineup and is framed as a flagship for general agentic work, complex software engineering, and long-horizon tasks. Its open weights are published under the XiaomiMiMo organization on Hugging Face, giving teams direct access to the same checkpoints that power hosted deployments. The model's positioning centers on sustained, multi-step execution: it is advertised as capable of independently completing professional tasks that would take human experts days or weeks, chaining together more than a thousand tool calls in the process. That emphasis on autonomy, rather than single-turn responsiveness, signals a design intent aimed at agent frameworks and orchestrated pipelines rather than casual chat use cases.

A defining technical characteristic is the the cataloged API limit context window, which lets the model ingest large codebases, lengthy documents, or extended interaction histories without aggressive truncation. The OpenRouter listing highlights leading placements on benchmarks tied to coding and agentic evaluation, including ClawEval, GDPVal, and SWE-bench Pro, framing these as evidence of strong real-world software engineering performance. For practitioners, this combination suggests strong fit for autonomous coding agents, research assistants that must reason across very long inputs, and multi-step workflows where maintaining continuity across thousands of tool invocations matters more than raw single-prompt fluency.

EmpirioLabs AImimo-v2-5-promimo

Quick Info

Powered by
Provider
EmpirioLabs AI
Model key
mimo-v2-5-pro
Release date
Apr 22, 2026
Last updated
Apr 22, 2026
Knowledge cutoff
2024-12
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$2.175
Output token cost
$4.35

Limits

Output tokens
128,000 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare MiMo V2.5 Pro pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about MiMo V2.5 Pro

EmpirioLabs AI

Coverage

A community developer report on NVIDIA's DGX Spark forum details running Xiaomi's MiMo-V2.5 310B MoE backbone with the new DFlash speculative decoder on a 2× DGX Spark (GB10) pair, achieving 22.3 tok/s in eager mode and up to 66.9 tok/s on structured JSON output depending on workload (27.6 prose, 45.4 code, 55.1 math). The post provides concrete deployment requirements: only the dflash/ folder (2.9 GB) is needed alongside the NVFP4 community quant (170 GB, TP=2 fits with room for 131K bf16 KV), a 5-layer qwen3-arch drafter with hidden 4096 and SWA-1024 cross-attending to backbone hidden states from layers [0, 11, 23, 35, 47], and vLL

EmpirioLabs AI

Coverage

The official Xiaomi MiMo Hugging Face model card for MiMo-V2.5-Pro describes the model as an open-source Mixture-of-Experts (MoE) language model with 1.02T total parameters and 42B active parameters, using a hybrid attention architecture and 3-layer Multi-Token Prediction (MTP). It supports up to 1M tokens of context l The model card details the hybrid attention design as interleaving Sliding Window Attention (SWA) and Global Attention (GA) at a 6:1 ratio with a 128 sliding window, reducing KV-cache storage by nearly 7x while maintaining long-context performance via learnable attention sink bias. The three MTP modules with dense FFNs

Videos about MiMo V2.5 Pro

More models around MiMo V2.5 Pro