Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Venice AI logo

Model details

MiMo-V2.5

MiMo-V2.5 is Xiaomi's native omnimodal mixture-of-experts model, designed as a single architecture that can absorb text, images, video, and audio without bolting on separate pipelines. It carries 310B total parameters with 15B active per token, sitting on a 48-layer backbone made of one dense layer and 47 MoE layers. The routing uses 256 experts with a top-8 selection, paired with a hybrid attention scheme that alternates sliding-window attention over a 128-token window with full attention at a 5:1 ratio. Dedicated encoders handle the non-text streams: a 729M vision encoder and a 261M audio encoder feed into the same backbone, and a three-layer Multi-Token Prediction head sits on top of the stack. The model ships in native FP8 block-wise e4m3 weights, which keeps the footprint manageable enough that community reports note it can run across a pair of high-end desktop accelerators. Open weights make the whole package available for self-hosting and fine-tuning rather than locking users into a single provider's stack.

MiMo-V2.5 inherits its backbone from the MiMo-V2-Flash line, and the broader V2.5 generation was tuned to push agentic capability and long-horizon coherence. On Xiaomi's published Coding Agent benchmark, MiMo-V2.5 lands at 56.1, and it reaches 71.8 on SWE-Bench Pro, putting it ahead of the previous V2-Pro generation and within striking distance of much larger frontier systems. The release notes emphasize complex software engineering and multi-step workflows as the primary intended use cases, and the model is wired for tool calling and structured reasoning out of the box. Hosting providers benchmark throughput north of 130 tokens per second, which positions MiMo-V2.5 as a practical option for production agents that need both multimodal grounding and fast iteration. For builders looking ahead, the combination of open weights, a compact active-parameter count, and native multimodal encoders makes it a flexible base for everything from autonomous coding assistants to video-and-audio-aware enterprise tools.

Venice AIxiaomi-mimo-v2-5mimo

Quick Info

Powered by
Provider
Venice AI
Model key
xiaomi-mimo-v2-5
Release date
Jun 11, 2026
Last updated
Jun 11, 2026
Knowledge cutoff
2024-12
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.40
Output token cost
$2.00

Limits

Output tokens
65,536 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare MiMo-V2.5 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about MiMo-V2.5

Venice AI

Coverage

Xiaomi MiMo-v2.5 Series API Permanent Price Reduction Up to 99 ... Google has announced that Gemini is processing 746T per week; Claude is ...

OpenRouter

CoverageBenchmark

A community-authored NVIDIA developer-forums thread (a3refaat, May 2026) documents a reproducible recipe for serving a third-party NVFP4 quantization of MiMo-V2.5 (lukealonso/MiMo-V2.5-NVFP4) on a 2× DGX Spark / GB10 cluster. The setup uses vLLM 0.21.1rc1.dev39 on CUDA 13.2 with tensor parallelism = 2, instanttensor lo Three concrete correctness fixes are detailed: (1) the ModelOpt mixed-precision path must dispatch MXFP8 linear layers to the MXFP8 method rather than falling through when the checkpoint mixes MXFP8 dense layers with NVFP4 experts; (2) the weight_scale_inv field is UE8M0 MXFP8 scale metadata and should not be reciproca

Venice AI

CoverageBenchmark

MiMo-V2.5 is a native omnimodal model by Xiaomi. $0.105 per million input tokens, $0.28 per million output tokens. 1,048,576 token context window. Higher uptime with 5 providers. Includes independent benchmarks from Artificial Analysis.

Videos about MiMo-V2.5

More models around MiMo-V2.5