Currently listed through these providers:
Model details
Gemini 3.5 Flash
Gemini 3.5 Flash sits in Google's Flash family as a high-efficiency multimodal model designed to deliver near-Pro reasoning and coding quality at substantially lower cost and latency. It is positioned around coding proficiency and parallel agentic execution loops, with a configurable thinking-effort dial that defaults to medium for fast responses while still allowing minimal, low, and high settings for finer cost-versus-performance trade-offs. The Inputs accept text, images, video, audio, and PDF documents, while outputs are plain text, making the model well suited to assistants that need to ingest mixed media and reason or write code in response.
Beyond raw chat, the model is tuned for production agentic workflows, including coding agents, sub-agent orchestration, and long-horizon tasks that chain many tool calls together, and it is generally available and marked stable for scaled production routing. Public pricing on Google's Vertex-backed routes is set at $1.50 per million input tokens and $9.00 per million output tokens, with cache reads available at a discounted rate, which together support cost-sensitive deployments that still need strong tool use. Practical fit therefore leans toward Copilot users who want responsive, multimodal coding help, lightweight agent loops, and integration with existing GitHub workflows without paying flagship-tier prices.
Quick Info
Powered by- Provider
- GitHub Copilot
- Model key
- gemini-3.5-flash
- Release date
- May 19, 2026
- Last updated
- May 19, 2026
- Knowledge cutoff
- 2025-01
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $1.50
- Output token cost
- $9.00
Limits
- Input tokens
- 128,000 tokens
- Output tokens
- 64,000 tokens
- Context window
- 200,000 tokens
Latest news about Gemini 3.5 Flash
Videos about Gemini 3.5 Flash
Recent tweets and retweets from GitHub Copilot
More models around Gemini 3.5 Flash
This exact model name is also listed by 22 other providers.