Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
CrossModel logo

Model details

Gemini 3.5 Flash

Gemini 3.5 Flash sits in the Flash tier of the Gemini family, positioned by OpenRouter as a high-efficiency multimodal model that brings near-Pro level coding and reasoning at Flash-tier cost and speed. It accepts text, image, video, audio, and PDF inputs, making it suitable for tasks that mix documents, media, and structured data in a single prompt. The model defaults to a medium thinking effort for faster, more cost-efficient responses, while still exposing configurable thinking levels — minimal, low, medium, and high — so teams can dial the reasoning budget to match the workload, from lightweight routing to deeper multi-step analysis.

In practical terms, the model is described as highly optimized for coding proficiency and for parallel agentic execution loops, which fits workflows such as code generation, tool-using agents, and orchestrated pipelines that need many smaller reasoning steps rather than a single long synthesis. A 1M-token context window lets it hold large codebases, long document collections, or extended conversation histories, while OpenRouter's provider table shows it served from Google Vertex with strong throughput. For builders, the combination of adjustable thinking, broad multimodal input, and agentic focus makes Gemini 3.5 Flash a fit when speed and cost matter but reasoning quality and tool use still need to feel close to a flagship model.

CrossModelgemini/gemini-3.5-flashgemini-flash

Quick Info

Powered by
Provider
CrossModel
Model key
gemini/gemini-3.5-flash
Release date
May 19, 2026
Last updated
May 19, 2026
Knowledge cutoff
2025-01
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$1.50
Output token cost
$9.00

Limits

Output tokens
65,536 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare Gemini 3.5 Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Gemini 3.5 Flash

Opper

CoverageRelease Notes

The official Gemini API documentation page for Gemini 3.5 Flash confirms that the model is generally available and stable for scaled production use, with the stable model ID gemini-3.5-flash. It lists a 1M token context window, 65k max output tokens, thinking support, and the same set of tools and platform features as The same page documents API-relevant changes for 3.5 Flash: a new default thinking effort level moved from high to medium, an improved low thinking mode that offers better quality at lower latency and cost for code and agentic tasks needing fewer steps, automatic thought preservation across multi-turn conversations wit

Opper

Coverage

Google Cloud CEO Thomas Kurian summarized Google I/O 2026 announcements on the Google Cloud blog, introducing the Gemini 3.5 model family as starting with Gemini 3.5 Flash and framing the launch as frontier intelligence combined with action. The post notes that Gemini 3.5 Flash is made available to Google Cloud custome The same Google Cloud post highlights supporting enterprise and developer infrastructure including eighth-generation TPUs used to train Gemini 3.5, the Managed Agents API on Agent Platform for building custom agents inside Google-hosted environments, CodeMender as an AI security agent on Agent Platform, Gemini Spark as

Opper

CoverageBenchmark

A Digital Applied technical guide dated May 19, 2026 confirms the Gemini 3.5 Flash GA details, listing the stable API model ID as gemini-3.5-flash and noting that it replaces the earlier gemini-3-flash-preview identifier used during the preview window. It reports an input context window of about 1.05M tokens, 65,536 ma The guide restates the vendor benchmark table from Google showing 3.5 Flash leading Claude Opus 4.7 and GPT-5.5 on five evaluations including MCP Atlas at 83.6 percent and CharXiv Reasoning at 84.2 percent, and covers the new thinking-level API surface along with migration guidance from gemini-3-flash-preview, anchorin

Opper

Coverage

Google DeepMind announced Gemini 3.5 Flash on May 19, 2026, kicking off the Gemini 3.5 family with a Flash-tier model aimed at agentic and coding workloads. According to the official Google blog post, 3.5 Flash is available the same day in the Gemini app, AI Mode in Google Search, the Google Antigravity agent-first dev The blog post cites vendor-reported benchmark wins including Terminal-Bench 2.1 at 76.2 percent, GDPval-AA at 1656 Elo, MCP Atlas at 83.6 percent, and CharXiv Reasoning at 84.2 percent for multimodal understanding, alongside a claim that 3.5 Flash is roughly 4 times faster in output tokens per second than other frontie

Videos about Gemini 3.5 Flash

More models around Gemini 3.5 Flash