Currently listed through these providers:
Model details
MiMo-V2.5-TTS
MiMo-V2.5-TTS is part of Xiaomi's MiMo-V2.5 speech model series, positioned as an end-to-end synthesis system that pairs TTS with a companion ASR model to form a full input-output voice chain. Rather than exposing low-level knobs like pitch, speed, or emotion intensity, the model accepts director-style natural language prompts—for instance, requesting a gentle 30-year-old female voice at a slightly faster pace with a hint of surprise—and renders the described performance directly. This language-driven control is the central design choice, aimed at content creators, short-video producers, and audiobook teams who want expressive output without manual parameter tuning.
The model's broader capability set centers on expressive, character-rich synthesis that goes beyond reading text aloud. It natively handles several Chinese dialects including Cantonese, Sichuanese, Henan, and Taiwanese accent, producing real dialect grammar and pronunciation rather than surface-level imitation. Within a single utterance it can switch emotions and sustain singing-style pitch, enabling performance-like delivery for narration, dialogue, and vocal content. The official release frames the V2.5 TTS series together with an ASR counterpart as a coordinated speech stack for the agent era, where machines both transcribe complex audio and shape voices on demand, making it well suited to projects that need flexible, controllable voice generation alongside recognition.
Quick Info
Powered by- Provider
- Xiaomi Token Plan (Europe)
- Model key
- mimo-v2.5-tts
- Release date
- Apr 22, 2026
- Last updated
- Apr 22, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 8,192 tokens
- Context window
- 8,192 tokens
Latest news about MiMo-V2.5-TTS
Videos about MiMo-V2.5-TTS
More models around MiMo-V2.5-TTS
This exact model name is also listed by 2 other providers.