Currently listed through these providers:
Model details
Gemini 3.5 Transcribe
Gemini 3.5 Transcribe is a specialized speech-to-text model introduced as part of the broader Gemini family, designed for precise and intelligent real-time transcription of audio. The model's announcement on the official Google blog frames it as an upgrade aimed at delivering more intelligent transcription behavior compared to prior speech-to-text offerings, positioning it as a focused tool rather than a general-purpose conversational model.
Through the unified Gemini API documentation surface, the model is presented as a dedicated transcription endpoint, making it well suited for developers who need reliable audio-to-text conversion in production pipelines. Its narrow specialization around speech recognition suggests a practical fit for use cases such as meeting transcription, captioning, voice note processing, and other workflows where faithful conversion of spoken content into text is the primary requirement.
Quick Info
Powered by- Provider
- Vercel AI Gateway
- Model key
- google/gemini-3.5-transcribe
- Release date
- Aug 26, 2026
- Last updated
- Aug 26, 2026
- Input modalities
- Output modalities
- Capabilities
- Base catalog fields only
Cost
- Input token cost
- $2.00
- Output token cost
- $12.00
Limits
- Output tokens
- 0 tokens
- Context window
- 0 tokens