Sulat.com
AI models
ZenMux logo

Model details

MiMo V2 Omni

MiMo-V2-Omni is an omni-modal foundation model introduced by Xiaomi on March 18, 2026, designed to bridge rich perception with autonomous action. Rather than treating vision, sound, and language as separate features, the model fuses dedicated image, video, and audio encoders into a single shared backbone so that it can see, hear, and read in one continuous stream. Training is framed around anticipating what comes next rather than only describing the present, meaning perception and agency are learned together from the first step rather than chained as two stages.

In practical terms, MiMo-V2-Omni is aimed at agentic workflows where a system must interpret messy, multimodal reality and then do something useful with it, such as a robotic arm responding to voice instructions, a digital agent executing multi-step software tasks, or an autonomous system fusing real-time sensor data. It natively supports structured tool calling, function execution, and UI grounding on the output side, and it combines visual grounding, multi-step planning, tool use, and code execution to handle tasks that span modalities. With an approximately 256K–262K token context window, the model is a strong fit for long, multi-document agentic pipelines where reasoning over combined text, image, video, and audio context must translate into reliable real-time action.

ZenMuxxiaomi/mimo-v2-omnimimo

Quick Info

Powered by
Provider
ZenMux
Model key
xiaomi/mimo-v2-omni
Release date
Mar 18, 2026
Last updated
Mar 18, 2026
Knowledge cutoff
2024-12
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.40
Output token cost
$2.00

Limits

Output tokens
265,000 tokens
Context window
265,000 tokens

Latest news about MiMo V2 Omni

Videos about MiMo V2 Omni

More models around MiMo V2 Omni