Sulat.com
AI models
Xiaomi Token Plan (China) logo

Model details

MiMo-V2-TTS

MiMo-V2-TTS is Xiaomi's in-house speech synthesis model introduced under the MiMo-V2 foundation model banner, a three-model lineup that also covers a large language model and a full-modality sibling. Independent reporting describes MiMo-V2-TTS as trained on hundreds of millions of hours of audio data, an unusually large audio corpus that signals a focus on broad vocal coverage and natural-sounding generation rather than a narrow task-specific voice. The model converts text input into spoken audio, fitting a role as the dedicated voice-generation piece of Xiaomi's broader MiMo-V2 stack alongside its reasoning and multimodal siblings.

Within the Xiaomi Token Plan ecosystem, MiMo-V2-TTS is positioned for developers building voice interfaces, narration, accessibility features, and multilingual spoken content who want a foundation-model-grade synthesizer rather than a small task-tuned TTS. Its placement in the same release wave as a trillion-parameter reasoning model and a full-modality model suggests Xiaomi intends it as a general-purpose speech backbone, suitable for conversational assistants, media production, and embedded device voice experiences. The open-weights availability on the platform lowers the barrier for local fine-tuning and integration, making it a practical choice when teams want control over voice style and deployment.

Xiaomi Token Plan (China)mimo-v2-ttsmimo

Quick Info

Powered by
Provider
Xiaomi Token Plan (China)
Model key
mimo-v2-tts
Release date
Mar 18, 2026
Last updated
Mar 18, 2026
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
8,192 tokens
Context window
8,192 tokens

Latest news about MiMo-V2-TTS

No articles yet. Fetch the latest news to show it here.

Videos about MiMo-V2-TTS

More models around MiMo-V2-TTS