Sulat.com
AI models
Xiaomi Token Plan (China) logo

Model details

MiMo-V2.5-TTS

MiMo-V2.5-TTS is Xiaomi MiMo's flagship neural text-to-speech family, designed to turn written input into natural, fluent audio with a high degree of stylistic control. The base model ships with a catalogue of out-of-the-box built-in voices that require no extra configuration, and the family extends this with companion variants for voice design from text descriptions and for voice replication from arbitrary audio samples. Across the series, users can shape output through controls for speed, emotion, role-play, and dialect, giving the system a flexible expressive range suited to narration, dialogue, and localised content. A low-latency streaming interface is available, returning audio in real time so the model fits naturally into interactive and conversational pipelines.

The V2.5 release marks the successor path for anyone still on the earlier V2 line, which the MiMo documentation marks as deprecated, steering users toward the V2.5 series as the actively supported option. Practically, the family suits teams that need a versatile speech engine covering quick prototyping with stock voices, custom character voices for games and media, and faithful cloning for branded assistants, all behind a consistent streaming API. The availability of the voice-clone variant through third-party inference hosting also signals broader ecosystem reach beyond Xiaomi's own platform, making it easier to integrate the same voice identity across multiple deployment environments.

Xiaomi Token Plan (China)mimo-v2.5-ttsmimo

Quick Info

Powered by
Provider
Xiaomi Token Plan (China)
Model key
mimo-v2.5-tts
Release date
Apr 22, 2026
Last updated
Apr 22, 2026
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
8,192 tokens
Context window
8,192 tokens

Latest news about MiMo-V2.5-TTS

Videos about MiMo-V2.5-TTS

More models around MiMo-V2.5-TTS