MiMo-V2-TTS is a speech synthesis model built to give AI agents a natural, expressive voice. Based on Xiaomi's self-developed Audio Tokenizer and a multi-codebook speech-text joint modeling architecture, the model enables granular control over tone, emotion, and speaking style. It can shift tone and emotional expression mid-sentence, mimicking the natural rhythm of human speech. The system intelligently reads text signals like punctuation and emphasis markers to generate appropriate speech without manual annotation. Supporting multiple Chinese dialects including Sichuanese, Cantonese, and Taiwanese, the model also handles singing with accurate pitch and rhythm, making it versatile for applications from conversational AI to creative content.
The model was pre-trained on hundreds of millions of hours of speech data and refined using multi-dimensional reinforcement learning to balance stability with expressive range. This scale of training enables the hyper-realistic emotional control and cross-regional adaptability the model demonstrates. Released as an open-weight model, MiMo-V2-TTS serves as a foundational tool for developers building expressive voice capabilities into agents, robots, and interactive products. Xiaomi positions the model as part of a broader voice pipeline for the agent era, with plans to expand multilingual coverage and deeper integration across its ecosystem.