Currently listed through these providers:
Model details
MiMo-V2.5-TTS-VoiceClone
MiMo-V2.5-TTS-VoiceClone is a specialized text-to-speech variant within Xiaomi's mimo family that focuses on voice cloning from short audio references. Its intended use is converting written text into natural, fluent speech while letting users configure parameters such as speaking style and target voice. According to the model's listing on DeepInfra, it sits in the Text To Speech category and is described as automatically converting input text into natural and fluent speech output, with the ability to precisely replicate voices from audio samples to enable speech synthesis of any voice.
The model extends the broader MiMo V2.5 family into speech generation, broadening Xiaomi's ecosystem beyond the text-focused general MiMo-V2.5 endpoint that supports multimodal input and very large context windows. For developers, this TTS variant is aimed at use cases such as audiobook narration, personalized voice assistants, and content localization where a consistent synthetic voice is needed. Being available as an open-weights release with no input or output token cost makes it accessible for experimentation and integration into both hobby projects and production pipelines that need voice-cloning capabilities without per-character billing.
Quick Info
Powered by- Provider
- Xiaomi Token Plan (Singapore)
- Model key
- mimo-v2.5-tts-voiceclone
- Release date
- Apr 22, 2026
- Last updated
- Apr 22, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 8,192 tokens
- Context window
- 8,192 tokens
Latest news about MiMo-V2.5-TTS-VoiceClone
Videos about MiMo-V2.5-TTS-VoiceClone
More models around MiMo-V2.5-TTS-VoiceClone
This exact model name is also listed by 2 other providers.