Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Xiaomi Token Plan (Europe) logo

Model details

MiMo-V2.5

MiMo-V2.5 is positioned as a major step toward agentic and multimodal understanding, built as a sparse mixture-of-experts model with 310B total parameters and 15B active, trained on 48 trillion tokens. Its language backbone inherits the hybrid sliding-window attention design from MiMo-V2-Flash, then gains native visual and audio understanding through in-house encoders joined by lightweight projectors. Training unfolds across five stages, from text pre-training and projector warmup through multimodal pre-training, supervised fine-tuning, agentic post-training that progressively widens the context window from 32K to 256K and onward to 1M tokens, and finally reinforcement learning plus an MOPD stage that strengthens perception, reasoning, and acting on what the model perceives.

On agentic benchmarks relevant to real deployment, MiMo-V2.5 delivers best-in-class scores against peer open models such as MiMo-V2-Pro, Kimi K2.6, and DeepSeek-V4-Flash, while also holding its own against frontier proprietary systems like Claude Opus 4.6, Gemini 3.1 Pro, and GPT-5.4. Because weights are openly available, the model can be self-hosted and quantized, and community runs have already demonstrated NVFP4 deployment on two DGX Spark units with vLLM, including full multimodal serving. The combination of long-context handling, native cross-modal reasoning, and open availability makes it a practical fit for teams building coding agents, document and media analysis tools, and other applications that need a single model that can see, hear, read, and act.

Xiaomi Token Plan (Europe)mimo-v2.5mimo

Quick Info

Powered by
Provider
Xiaomi Token Plan (Europe)
Model key
mimo-v2.5
Release date
Apr 22, 2026
Last updated
Apr 22, 2026
Knowledge cutoff
2024-12
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
131,072 tokens
Context window
1,048,576 tokens

Latest news about MiMo-V2.5

Xiaomi Token Plan (Europe)

CoverageBenchmark

Hey guys, I decided to stop lurking and make an effort to contribute in this forum. I got lukealonso/MiMo-V2.5-NVFP4 running on a 2× DGX Spark / GB10 cluster with vLLM, including Omni/multimodal serving and MTP speculat…

Xiaomi Token Plan (Europe)

CoverageBenchmark

A community user (a3refaat) posted a detailed recipe on the NVIDIA DGX Spark / GB10 forum for running the third-party quantized checkpoint lukealonso/MiMo-V2.5-NVFP4 on a 2× DGX Spark cluster with TP=2. The setup uses vLLM 0.21.1rc1.dev39 (CUDA 13.2), instanttensor load format, Triton attention with diffkv, FlashInfer- Three substantive correctness fixes are highlighted alongside the recipe: (1) the ModelOpt mixed-precision path must dispatch MXFP8 linear layers to the MXFP8 method rather than fall through when the checkpoint mixes MXFP8 dense layers with NVFP4 experts, (2) weight_scale_inv should be treated as UE8M0 MXFP8 scale meta

Videos about MiMo-V2.5

More models around MiMo-V2.5