Currently listed through these providers:
Model details
Voxtral Mini TTS (latest)
Voxtral Mini TTS latest is Mistral's compact text-to-speech endpoint, designed for turning written prompts into spoken audio with a small, predictable footprint. The standout capability is zero-shot voice cloning, which lets users generate speech in a target voice from a brief sample rather than requiring a fully trained speaker profile, paired with multilingual synthesis so the same pipeline can serve content across languages without swapping models. A 4K context window is documented on the Opper gateway listing, giving enough room for reasonably long scripts while keeping the interaction snappy and focused on speech rather than long-form reasoning.
In practical terms, the model fits workflows where natural-sounding voiceover, narration, dubbing, or accessibility audio needs to be produced on demand from text input, and the zero-shot cloning approach makes it attractive for prototyping branded voices or personalized assistants without bespoke training runs. Pricing is billed per character rather than per token, which simplifies budgeting for scripted content like audiobook chapters, e-learning narration, or in-app voice prompts, and the EU residency option with an Enterprise-tier zero data retention posture and available GDPR DPA make it viable for European deployments with stricter data-handling requirements. The versioned snapshot slug "voxtral-mini-tts-2603" indicates active snapshotting under the "latest" alias, suggesting users can expect rolling updates as Mistral refines the underlying speech synthesis quality.
Quick Info
Powered by- Provider
- Mistral
- Model key
- voxtral-mini-tts-latest
- Release date
- Mar 1, 2026
- Last updated
- Mar 1, 2026
- Input modalities
- Output modalities
- Capabilities
- Base catalog fields only
Limits
- Output tokens
- 0 tokens
- Context window
- 0 tokens