Currently listed through these providers:
Model details
Gemini 2.5 Flash TTS
Gemini 2.5 Flash TTS Preview is Google's text-to-speech variant built from the Gemini 2.5 Flash lineage, focused on converting written input into synthesized audio for use cases such as voice assistants, narration, accessibility, and content production pipelines. Third-party aggregator CloudPrice classifies it as an active "Text to Speech" model created by Google, and Google's own documentation exposes the preview under the Gemini API developer docs, signaling first-party support rather than a community experiment. The preview framing suggests the model is intended to evolve alongside the broader Gemini 2.5 family as audio generation matures inside the Gemini ecosystem.
In practice, the model is positioned as a lightweight, speech-first endpoint rather than a general-purpose multimodal assistant: it accepts text as input and produces audio as output, making it well suited for workflows that need natural-sounding voice from written prompts without invoking a larger reasoning model. CloudPrice documents a working envelope of an 8K-token context window and up to 16K output tokens of synthesized speech, giving integrators room for long prompts or extended utterances. The preview designation makes it a reasonable choice for teams prototyping conversational interfaces, narration layers, or accessibility features who want to evaluate Flash-tier latency and cost before committing to a generally available release.
Quick Info
Powered by- Provider
- Vertex
- Model key
- gemini-2.5-flash-tts
- Release date
- Sep 30, 2025
- Last updated
- Dec 10, 2025
- Knowledge cutoff
- 2025-01
- Input modalities
- Output modalities
- Capabilities
- Base catalog fields only
Cost
- Input token cost
- $0.50
- Output token cost
- $10.00
Limits
- Output tokens
- 16,384 tokens
- Context window
- 32,768 tokens
Latest news about Gemini 2.5 Flash TTS
No articles yet. Fetch the latest news to show it here.