CloudPrice catalogs Gemini 3.1 Flash TTS Preview as a Google text-to-speech model with an 8K context window, up to 16K output tokens, text input and audio output modalities, reasoning effort default, and a knowledge cutoff listed as December 2024. The page marks the model as Active and Current, and identifies the canon A version table on the same page enumerates sibling Gemini Flash releases ranging from Gemini 3 Flash in December 2025 through later variants dated through September 2026, placing Gemini 3.1 Flash TTS Preview alongside them. The page does not trace its specifications or pricing to a primary Google source, and the conte
Model details
Gemini 3.1 Flash TTS Preview
Gemini 3.1 Flash TTS Preview is a Google text-to-speech model positioned as a preview voice variant of the Gemini 3.1 Flash line, aimed at giving early access to fast voice synthesis. It is listed as an Active model on the Google Gemini provider with an 8K-token context window and support for up to 16K output tokens, giving room for longer transcripts within a single generation call. The framing reflects a lineage strategy where the Flash-class model is reused and specialized for audio output rather than introducing a separate architecture from scratch.
Out of the box, the model natively interprets a transcript and decides how words should sound, producing natural delivery without any extra prompting. When more control is needed, it can be directed through natural-language instructions that describe an audio profile (who is speaking and what their voice sounds like), the scene or environment, and any extra director's notes that shape the performance. Inline tags such as "whispers", "laughs", "cough", "sighs", or "gasp" can be placed within the transcript to fine-tune the tone, pace, and emotion of specific lines, making the model a flexible fit for narration, character voice work, and guided voice-over production.
Quick Info
Powered by- Provider
- Model key
- gemini-3.1-flash-tts-preview
- Release date
- Apr 15, 2026
- Last updated
- Apr 15, 2026
- Knowledge cutoff
- 2025-01
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $1.00
- Output token cost
- $20.00
Limits
- Output tokens
- 16,384 tokens
- Context window
- 8,192 tokens
Latest news about Gemini 3.1 Flash TTS Preview
BenchLM tracks Gemini 3.1 Flash TTS Preview as a released proprietary, non-reasoning text-to-speech model with a recorded release date of April 13, 2026, a 32K context window, text input and audio output modalities, and a preview lifecycle. The profile lists the API model ID as gemini-3.1-flash-tts-preview and cites Go An owner-defined Audio Realism Benchmark section reports an Elo-style rating of 1,046, a 39.0% win rate across 475 battles, and an average generation time of 4.51 seconds from an owner snapshot dated August 31, 2026. Spec fields including maximum output, knowledge cutoff, parameter count, cloud regions, rate limits, an