Currently listed through these providers:
Model details
Gemini 3.5 Flash
Gemini 3.5 Flash sits in Google's Gemini Flash family and is positioned as a lightweight, high-throughput variant of the Gemini 3.5 generation, engineered for latency-sensitive, high-volume deployments rather than the deepest reasoning tier. Independent tracking notes the model reached general availability and stable production status in May 2026, with the official model ID gemini-3.5-flash confirmed against Google's Gemini API documentation. That placement in the Flash lineup signals a design emphasis on speed and cost efficiency, while still inheriting the multimodal capabilities that the broader Gemini family is known for, making it a natural fit for applications that need to balance responsiveness with strong general capability.
In practical terms, the model is shaped for workloads where quick, reliable responses matter more than maximum depth, such as interactive assistants, routing layers in agentic pipelines, code generation helpers, and retrieval-augmented chat where long inputs are common. Its very large context window lets teams pass substantial documents, transcripts, or tool histories in a single request, which is useful for sub-agent orchestration and multi-step workflows that previously required chunking. As part of the Flash tier, it serves as an accessible default that teams can route the bulk of traffic to, reserving heavier Gemini variants for cases that demand extra reasoning or longer, more elaborate outputs.
Quick Info
Powered by- Provider
- UnoRouter
- Model key
- gemini-3.5-flash
- Release date
- May 19, 2026
- Last updated
- May 19, 2026
- Knowledge cutoff
- 2025-01
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.1857
- Output token cost
- $1.1142
Limits
- Output tokens
- 65,536 tokens
- Context window
- 1,048,576 tokens
Latest news about Gemini 3.5 Flash
Videos about Gemini 3.5 Flash
More models around Gemini 3.5 Flash
This exact model name is also listed by 21 other providers.