Currently listed through these providers:
Model details
MiMo-V2-TTS
MiMo-V2-TTS is Xiaomi's in-house speech synthesis model introduced under the MiMo-V2 foundation model banner, a three-model lineup that also covers a large language model and a full-modality sibling. Independent reporting describes MiMo-V2-TTS as trained on hundreds of millions of hours of audio data, an unusually large audio corpus that signals a focus on broad vocal coverage and natural-sounding generation rather than a narrow task-specific voice. The model converts text input into spoken audio, fitting a role as the dedicated voice-generation piece of Xiaomi's broader MiMo-V2 stack alongside its reasoning and multimodal siblings.
Within the Xiaomi Token Plan ecosystem, MiMo-V2-TTS is positioned for developers building voice interfaces, narration, accessibility features, and multilingual spoken content who want a foundation-model-grade synthesizer rather than a small task-tuned TTS. Its placement in the same release wave as a trillion-parameter reasoning model and a full-modality model suggests Xiaomi intends it as a general-purpose speech backbone, suitable for conversational assistants, media production, and embedded device voice experiences. The open-weights availability on the platform lowers the barrier for local fine-tuning and integration, making it a practical choice when teams want control over voice style and deployment.
Quick Info
Powered by- Provider
- Xiaomi Token Plan (China)
- Model key
- mimo-v2-tts
- Release date
- Mar 18, 2026
- Last updated
- Mar 18, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 8,192 tokens
- Context window
- 8,192 tokens
Latest news about MiMo-V2-TTS
No articles yet. Fetch the latest news to show it here.
Videos about MiMo-V2-TTS
More models around MiMo-V2-TTS
This exact model name is also listed by 2 other providers.