Currently listed through these providers:
Model details
Gemini 3.5 Flash
Gemini 3.5 Flash continues the Gemini 3 series as a natively multimodal reasoning model, extending the Gemini 3 Flash foundation with configurable thinking levels that let developers tune the balance of quality, cost, and latency for each workload. The DeepMind model card positions it as the iteration that combines frontier intelligence with action, aligning with the broader Google I/O 26 push toward agentic enterprise tooling and the Gemini Enterprise Agent Platform announced on Google Cloud. Inputs are described as a mix of text strings such as questions, prompts, or documents to summarize, along with images, audio, and video files, keeping it versatile for grounded, real-world tasks.
Practically, Gemini 3.5 Flash fits scenarios that need quick, multimodal understanding without sacrificing reasoning depth, from summarizing mixed-media documents to powering agents that take action across enterprise systems. Its lineage from the Gemini 3 Flash reasoning base means teams familiar with prior Gemini Flash behavior can expect continuity in instruction-following while gaining finer control over how much the model thinks before responding. Combined with its release alongside Google Cloud's eighth-generation TPUs and the Agentic Data Cloud, it is well suited for production pipelines where latency budgets are tight but multimodal comprehension and tool-driven reasoning still matter.
Quick Info
Powered by- Provider
- Vertex
- Model key
- gemini-3.5-flash
- Release date
- May 19, 2026
- Last updated
- May 19, 2026
- Knowledge cutoff
- 2025-01
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $1.50
- Output token cost
- $9.00
Limits
- Output tokens
- 65,536 tokens
- Context window
- 1,048,576 tokens
Latest news about Gemini 3.5 Flash
Videos about Gemini 3.5 Flash
Recent tweets and retweets from Vertex
More models around Gemini 3.5 Flash
This exact model name is also listed by 22 other providers.