Sulat.com
AI models
Xiaomi Token Plan (Europe) logo

Model details

MiMo-V2.5-TTS

MiMo-V2.5-TTS is part of Xiaomi's MiMo-V2.5 speech model series, positioned as an end-to-end synthesis system that pairs TTS with a companion ASR model to form a full input-output voice chain. Rather than exposing low-level knobs like pitch, speed, or emotion intensity, the model accepts director-style natural language prompts—for instance, requesting a gentle 30-year-old female voice at a slightly faster pace with a hint of surprise—and renders the described performance directly. This language-driven control is the central design choice, aimed at content creators, short-video producers, and audiobook teams who want expressive output without manual parameter tuning.

The model's broader capability set centers on expressive, character-rich synthesis that goes beyond reading text aloud. It natively handles several Chinese dialects including Cantonese, Sichuanese, Henan, and Taiwanese accent, producing real dialect grammar and pronunciation rather than surface-level imitation. Within a single utterance it can switch emotions and sustain singing-style pitch, enabling performance-like delivery for narration, dialogue, and vocal content. The official release frames the V2.5 TTS series together with an ASR counterpart as a coordinated speech stack for the agent era, where machines both transcribe complex audio and shape voices on demand, making it well suited to projects that need flexible, controllable voice generation alongside recognition.

Xiaomi Token Plan (Europe)mimo-v2.5-ttsmimo

Quick Info

Powered by
Provider
Xiaomi Token Plan (Europe)
Model key
mimo-v2.5-tts
Release date
Apr 22, 2026
Last updated
Apr 22, 2026
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
8,192 tokens
Context window
8,192 tokens

Latest news about MiMo-V2.5-TTS

Videos about MiMo-V2.5-TTS

More models around MiMo-V2.5-TTS