Sulat.com
AI models
Xiaomi Token Plan (China) logo

Model details

MiMo-V2.5-TTS-VoiceDesign

MiMo-V2.5-TTS-VoiceDesign is Xiaomi's text-to-speech entry focused on generative voice creation rather than fixed preset speakers. According to Xiaomi's official MiMo documentation, the VoiceDesign variant is tailored for synthesizing speech purely from text descriptions, automatically generating a voice without requiring preset speakers or reference audio. This distinguishes it from sibling offerings in the same series, where a separate voice clone model is reserved for replicating arbitrary speakers from audio samples and the base TTS model relies on built-in voices. The model sits within the broader MiMo-V2.5-TTS family, which has been positioned as the active lineage after the prior V2 series was deprecated and users were directed to migrate to V2.5.

The VoiceDesign model's defining capability is text-prompt-driven voice persona generation, supplemented by fine-grained control over delivery style. The MiMo docs describe the series as supporting adjustable speed, emotion, role-play, and dialect rendering, with low-latency streaming output restored so the interface returns responses in real time for the standard TTS variant. Input is restricted to text while output is audio, and third-party catalog listings show an Active status under the canonical identifier xiaomi-mimo-2-5-tts-voicedesign. Practically, this makes the model a strong fit for applications that need novel, on-demand voices without audio references, such as character dialogue, branded narration, or localized storytelling, while teams needing exact voice replication should pair it with the voice clone sibling rather than use VoiceDesign in isolation.

Xiaomi Token Plan (China)mimo-v2.5-tts-voicedesignmimo

Quick Info

Powered by
Provider
Xiaomi Token Plan (China)
Model key
mimo-v2.5-tts-voicedesign
Release date
Apr 22, 2026
Last updated
Apr 22, 2026
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
8,192 tokens
Context window
8,192 tokens

Latest news about MiMo-V2.5-TTS-VoiceDesign

Videos about MiMo-V2.5-TTS-VoiceDesign

More models around MiMo-V2.5-TTS-VoiceDesign