Sulat.com
AI models
Xiaomi Token Plan (Europe) logo

Model details

MiMo-V2-TTS

MiMo-V2-TTS is a speech synthesis model built to give AI agents a natural, expressive voice. Based on Xiaomi's self-developed Audio Tokenizer and a multi-codebook speech-text joint modeling architecture, the model enables granular control over tone, emotion, and speaking style. It can shift tone and emotional expression mid-sentence, mimicking the natural rhythm of human speech. The system intelligently reads text signals like punctuation and emphasis markers to generate appropriate speech without manual annotation. Supporting multiple Chinese dialects including Sichuanese, Cantonese, and Taiwanese, the model also handles singing with accurate pitch and rhythm, making it versatile for applications from conversational AI to creative content.

The model was pre-trained on hundreds of millions of hours of speech data and refined using multi-dimensional reinforcement learning to balance stability with expressive range. This scale of training enables the hyper-realistic emotional control and cross-regional adaptability the model demonstrates. Released as an open-weight model, MiMo-V2-TTS serves as a foundational tool for developers building expressive voice capabilities into agents, robots, and interactive products. Xiaomi positions the model as part of a broader voice pipeline for the agent era, with plans to expand multilingual coverage and deeper integration across its ecosystem.

Xiaomi Token Plan (Europe)mimo-v2-ttsmimo

Quick Info

Powered by
Provider
Xiaomi Token Plan (Europe)
Model key
mimo-v2-tts
Release date
Mar 18, 2026
Last updated
Mar 18, 2026
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
8,192 tokens
Context window
8,192 tokens

Latest news about MiMo-V2-TTS

Xiaomi Token Plan (Europe)

Coverage

Back in March, Xiaomi introduced its MiMo-V2-TTS speech synthesis model, which focuses on detailed control over tone, emotion, and speaking style. The company said at the time that it could handle everything from natural conversations to singing, with support for multiple Chinese dialects. Now, Xiaomi is updating that

Xiaomi Token Plan (Europe)

Coverage

Back in March, Xiaomi introduced its MiMo-V2-TTS speech synthesis model, which focuses on detailed control over tone, emotion, and speaking style. The company said at the time that it could handle everything from natural conversations to singing, with support for multiple Chinese dialects. Now, Xiaomi is updating that

Xiaomi Token Plan (Europe)

CoverageRelease Notes

Xiaomi unveiled MiMo-V2-Pro, MiMo-V2-Omni and MiMo-V2-TTS models on March 19, 2026, targeting AI agents, multimodal processing and speech synthesis12.

Xiaomi Token Plan (Europe)

Coverage

The MiMo-V2 lineup includes three models: MiMo-V2-Pro, MiMo-V2-Omni, and MiMo-V2-TTS.

Xiaomi Token Plan (Europe)

CoverageAnalysis

Analysis of the MiMo-V2-TTS model by Xiaomi and comparison to other Text to Speech models across key metrics including quality ELO, speed, and price.

Videos about MiMo-V2-TTS

More models around MiMo-V2-TTS