Sulat.com
AI models
Xiaomi Token Plan (Singapore) logo

Model details

MiMo-V2.5-TTS-VoiceClone

MiMo-V2.5-TTS-VoiceClone is a specialized text-to-speech variant within Xiaomi's mimo family that focuses on voice cloning from short audio references. Its intended use is converting written text into natural, fluent speech while letting users configure parameters such as speaking style and target voice. According to the model's listing on DeepInfra, it sits in the Text To Speech category and is described as automatically converting input text into natural and fluent speech output, with the ability to precisely replicate voices from audio samples to enable speech synthesis of any voice.

The model extends the broader MiMo V2.5 family into speech generation, broadening Xiaomi's ecosystem beyond the text-focused general MiMo-V2.5 endpoint that supports multimodal input and very large context windows. For developers, this TTS variant is aimed at use cases such as audiobook narration, personalized voice assistants, and content localization where a consistent synthetic voice is needed. Being available as an open-weights release with no input or output token cost makes it accessible for experimentation and integration into both hobby projects and production pipelines that need voice-cloning capabilities without per-character billing.

Xiaomi Token Plan (Singapore)mimo-v2.5-tts-voiceclonemimo

Quick Info

Powered by
Provider
Xiaomi Token Plan (Singapore)
Model key
mimo-v2.5-tts-voiceclone
Release date
Apr 22, 2026
Last updated
Apr 22, 2026
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
8,192 tokens
Context window
8,192 tokens

Latest news about MiMo-V2.5-TTS-VoiceClone

Videos about MiMo-V2.5-TTS-VoiceClone

More models around MiMo-V2.5-TTS-VoiceClone