Currently listed through these providers:
Model details
Gemini 3.1 Flash Live Preview
Gemini 3.1 Flash Live Preview is Google's answer to the growing demand for conversational AI that can listen, think, and speak in near real time. It is delivered through the Gemini Live API and is positioned as an ultra-low-latency, audio-to-audio model that combines high-fidelity multimodal reasoning with a streaming voice interface. The design intent is clear from the way Google split its documentation between a dedicated model page and a Live API overview: this release is meant for builders assembling voice agents, live customer-service assistants, and other interactive experiences where response delay, voice naturalness, and on-the-fly tool use matter more than raw text generation throughput.
In practical terms, the model gives developers a large working memory for long, multi-turn conversations while still allowing streamed audio output, tool calling, and image plus video understanding alongside text and speech input. Reported benchmark results underscore the strength of the underlying reasoning stack, with a 94% score on GPQA graduate-level science questions and a 44% score on the HLE expert reasoning benchmark, suggesting solid general knowledge and disciplined multi-step thinking behind the conversational surface. Migration teams coming from earlier native-audio previews should treat this as a fresh starting point rather than a drop-in swap, since Google also revised the thinking configuration, server event shape, and tool-use behavior alongside the launch.
Quick Info
Powered by- Provider
- Model key
- gemini-3.1-flash-live-preview
- Release date
- Mar 26, 2026
- Last updated
- Mar 26, 2026
- Knowledge cutoff
- 2025-01
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.75
- Output token cost
- $4.50
Limits
- Output tokens
- 65,536 tokens
- Context window
- 131,072 tokens
Latest news about Gemini 3.1 Flash Live Preview
No articles yet. Fetch the latest news to show it here.