Currently listed through these providers:
Model details
Gemini 3.5 Transcribe Live
Gemini 3.5 Transcribe Live is a dedicated speech-to-text model designed for continuous, real-time transcription through Google's Gemini Live API. The official Google AI for Developers documentation describes it as supporting low-latency transcription, and developers connect to it over WebSockets or through the Google Gen AI SDK to stream audio and receive incremental text as speech is detected. The model's identifier in the documentation is linked to a dedicated page under the Gemini API model reference, framing it as a specialized transcribe variant within the broader Gemini family rather than a general-purpose conversational model.
In practice, the model fits use cases that require streaming audio input and immediate text output, such as live captioning, meeting transcription, and voice-driven product interfaces. Because the Live API exposes it as a streaming endpoint rather than a batch transcription job, it suits applications where latency and incremental partial results matter more than offline analysis. Third-party coverage has highlighted headline figures around sub-half-second response times and competitive word error rates, pointing to a model positioned against other modern streaming ASR systems for production voice workflows.
Quick Info
Powered by- Provider
- Vercel AI Gateway
- Model key
- google/gemini-3.5-transcribe-live
- Release date
- Aug 26, 2026
- Last updated
- Aug 26, 2026
- Input modalities
- Output modalities
- Capabilities
- Base catalog fields only
Limits
- Output tokens
- 0 tokens
- Context window
- 0 tokens