Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
LLM Gateway logo

Model details

Gemini 3.7 Flash (Google Vertex AI)

Gemini 3.7 Flash is positioned by Google DeepMind as the most capable workhorse model yet in the Flash family, aimed squarely at complex agentic tasks that need to scale, especially coding workflows and multi-step tool-driven assistants. The official Flash model page introduces the release with the tagline "Best for tackling complex agentic tasks at scale" and the descriptor "Our most intelligent workhorse model yet for coding and agents," signalling an emphasis on reasoning depth, tool use, and orchestration rather than pure chat. Entry points exposed on the same page route developers into AI Studio with the model identifier gemini-3.7-flash, and consumers into the Gemini app, indicating the same weights serve both builder and end-user surfaces.

The DeepMind Flash page organizes the 3.7 Flash story around five reader journeys — Capabilities, Hands-on, Showcase, Performance, and Model information — and the catalog places 3.7 Flash in the gemini-flash family with very large context, structured output, temperature control, attachment handling, reasoning, and tool calling enabled. For practitioners, the practical fit is a fast, multimodal Flash-tier model intended to keep latency low while running long-horizon agent loops, code-generation pipelines, and tool-using assistants that benefit from a long context window. Within the Flash lineup it sits as the newer, agent-focused successor to earlier Gemini 3.x Flash variants, trading peak benchmark headline numbers for predictable throughput on production agent stacks.

LLM Gatewaygoogle-vertex/gemini-3.7-flashgemini-flash

Quick Info

Powered by
Provider
LLM Gateway
Model key
google-vertex/gemini-3.7-flash
Release date
Aug 13, 2026
Last updated
Aug 13, 2026
Knowledge cutoff
2026-03
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.75
Output token cost
$3.75

Limits

Output tokens
65,536 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare Gemini 3.7 Flash (Google Vertex AI) pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Gemini 3.7 Flash (Google Vertex AI)

LLM Gateway

Coverage

Google's official August 2026 AI recap explicitly names the launch of Gemini 3.7 Flash alongside Gemini 3.5 Transcribe and the Pixel 11 lineup. The post frames the model as a cost-efficient developer option, positioning it within Google's broader push to make AI more practical and accessible. The Gemini app also crossed one billion monthly users during the same period. Beyond the 3.7 Flash launch, the recap highlights expanded Gemini Live productivity tools, hands-free voice features across Google Workspace, and the rollout of Gemini in Chrome on Android. The post ties these launches to Google's stated goal of bringing intelligence directly into where people already work, signaling 3.7 Flash's role in the developer tier.

LLM Gateway

Coverage

On September 1, 2026, Google published a first-party blog post announcing agentic video understanding for Gemini, with Gemini 3.7 Flash explicitly named as one of the models receiving the new capability (alongside 3.6 Flash and 3.5 Flash-Lite). The post is authored by named Google DeepMind staff — Senior Product Manage The post reports concrete technical metrics for the agentic video path on Gemini 3.7 Flash: up to 88% reduction in token consumption, up to 66% reduction in analysis cost, and up to 7% improvement in accuracy across standard video analysis benchmarks. The capability is available today for video uploads and YouTube URLs

LLM Gateway

Coverage

Independent benchmark firm Artificial Analysis measured Gemini 3.7 Flash at 340.1 tokens per second in output and ranked it first among 186 models tested for generation speed. The figure comes from a Google AI Studio post announcing the model on August 13, 2026, just three weeks after Gemini 3.6 Flash. The article fram Beyond the headline throughput number, the article situates Gemini 3.7 Flash against competing fast tiers and highlights that Google released it specifically as a response to developer feedback and algorithmic innovations intended for future models. The blog does not distinguish between the Gemini API and Google Vertex

LLM Gateway

Coverage

SiliconANGLE's August 13, 2026 article, headlined "Google launches Gemini 3.7 Flash for coding, AI agent projects," provides independent tech-press confirmation of the model's launch date and intended use cases. The headline alone establishes that Google positioned Gemini 3.7 Flash for coding workflows and AI agent dev The supplied excerpt is dominated by site navigation and advertising chrome, so detailed benchmark scores, pricing, context window, or Vertex-AI-specific availability information are not visible in the captured text. The candidate therefore functions as independent corroboration of the launch event, creator (Google), d

LLM Gateway

Coverage

AIToolly reports that on August 13, 2026, Google DeepMind officially announced Gemini 3.7 Flash via the DeepMind Blog, marking the latest iteration in the Gemini "Flash" category optimized for speed and efficiency. The article cites the DeepMind Blog as the primary source and notes the release continues the Flash linea The piece explicitly attributes the model to Google DeepMind as creator and confirms the August 13, 2026 announcement date. Technical details beyond the launch confirmation, family positioning, and versioning rationale are not elaborated in the supplied excerpt, and the page itself is a third-party aggregator summarizi

Videos about Gemini 3.7 Flash (Google Vertex AI)

More models around Gemini 3.7 Flash (Google Vertex AI)