Currently listed through these providers:
Model details
MiMo V2 Omni
MiMo-V2-Omni is an omni-modal foundation model introduced by Xiaomi on March 18, 2026, designed to bridge rich perception with autonomous action. Rather than treating vision, sound, and language as separate features, the model fuses dedicated image, video, and audio encoders into a single shared backbone so that it can see, hear, and read in one continuous stream. Training is framed around anticipating what comes next rather than only describing the present, meaning perception and agency are learned together from the first step rather than chained as two stages.
In practical terms, MiMo-V2-Omni is aimed at agentic workflows where a system must interpret messy, multimodal reality and then do something useful with it, such as a robotic arm responding to voice instructions, a digital agent executing multi-step software tasks, or an autonomous system fusing real-time sensor data. It natively supports structured tool calling, function execution, and UI grounding on the output side, and it combines visual grounding, multi-step planning, tool use, and code execution to handle tasks that span modalities. With an approximately 256K–262K token context window, the model is a strong fit for long, multi-document agentic pipelines where reasoning over combined text, image, video, and audio context must translate into reliable real-time action.
Quick Info
Powered by- Provider
- ZenMux
- Model key
- xiaomi/mimo-v2-omni
- Release date
- Mar 18, 2026
- Last updated
- Mar 18, 2026
- Knowledge cutoff
- 2024-12
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.40
- Output token cost
- $2.00
Limits
- Output tokens
- 265,000 tokens
- Context window
- 265,000 tokens