Sulat.com
AI models
Requesty logo

Model details

Gemini 3.8 Flash

Gemini 3.8 Flash sits within the Gemini 3 family as a successor to Gemini 3.7 Flash, positioned by Google DeepMind as the most intelligent workhorse model in the lineup yet for coding and agents. According to the official model card, it brings performance advancements across software engineering and agentic knowledge workflows, retaining support for customizable effort levels that let teams tune the balance between quality, cost, and latency. The accompanying product page frames it as the right choice for tackling complex agentic tasks at scale, where Flash-level speed and latency still matter but the reasoning load is heavier than typical lightweight calls.

Beyond its positioning, Gemini 3.8 Flash benefits from a deliberate Google ecosystem rollout that includes both consumer-facing entry points and developer tooling. Google DeepMind exposes the model through the Gemini app and through AI Studio using the model identifier gemini-3.8-flash, giving practitioners a direct path to experiment, prototype, and integrate. The official pages organize the story around dedicated sections for capabilities, hands-on examples, showcase use cases, performance discussion, and model information, signaling that the release is meant to be evaluated not only on raw benchmarks but also on practical agentic and coding workflows. For teams building automated pipelines or coding assistants that need strong reasoning without stepping up to a larger flagship tier, Gemini 3.8 Flash offers a focused option in the mid-latency Flash segment.

Requestygemini-3.8-flashgemini-flash

Quick Info

Powered by
Provider
Requesty
Model key
gemini-3.8-flash
Release date
Sep 2, 2026
Last updated
Sep 2, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.75
Output token cost
$3.75

Limits

Output tokens
65,535 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare Gemini 3.8 Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Gemini 3.8 Flash

Requesty

CoverageBenchmark

DataCamp's coverage of Gemini 3.8 Flash and the Gemini 3.8 Flash Cyber variant frames the September 2, 2026 release as Google's third Flash model in six weeks and its strongest reasoning and coding entry in the Flash tier, arriving at the same speed and cost as 3.7 Flash three weeks earlier. Reported headline numbers i The piece flags uneven results: coding and tool use jumped, but Humanity's Last Exam stayed flat at 45.4% versus 45.7% for the prior Flash. Pricing is held at $0.75 input and $3.75 output per 1M tokens through December 31, 2026, then reverts to $1.50/$7.50. Gemini 3.8 Flash Cyber is restricted to trusted defenders via

Requesty

Official sourceBenchmark

The Vertex AI EU endpoint page on Requesty documents the regional deployment as `vertex/gemini-3.8-flash@eu`, served from the EU under Vertex AI Data Governance with no data retention and no training use. It reports a 1.0M-token context window, 66K-token max output, chat API type, and a price of $0.83 input and $4.13 o Workload cost examples are $0.12 for 100K input plus 10K output, $1.24 for 1M input plus 100K output, and $12.38 for 10M input plus 1M output. The page explicitly states no benchmarks are published for this exact regional variant and recommends the base model page for shared leaderboard scores, while integration uses t

Requesty

CoverageBenchmark

Coursiv's coverage dates the Gemini 3.8 Flash release to September 2, 2026, three weeks after Gemini 3.7 Flash and one day after Anthropic shipped Claude Fable 5.1. It describes the model as positioned for frontier-class results on coding, finance, and legal agent benchmarks at a fraction of rival pricing, with promoti Specs listed include a 1M-token context window, 64K-token maximum output, multimodal text/image/audio/video inputs, customizable effort levels trading quality against cost and latency, and a knowledge cutoff of March 2026 for some domains and January 2025 for others. Distribution at launch covers the Gemini app for AI

Requesty

Official sourceBenchmark

Requesty's model page for Gemini 3.8 Flash confirms Google's latest Flash-tier model is now routed through the Requesty aggregator under a single pinned ID, with the router automatically selecting across serving providers and failing over on price or health. The page reports a 1.0M-token context window, 66K-token max o Two Vertex AI endpoints are exposed: a global region at $0.75 input / $3.75 output per 1M tokens (list $1.50/$7.50) with cache reads at $0.07/1M, and an EU region at $0.83/$4.13 per 1M (list $1.65/$8.25) with cache reads at $0.08/1M. Requesty-reported benchmark numbers include a 76.3% Coding Index, 95.3% on GPQA Diamon

Videos about Gemini 3.8 Flash

More models around Gemini 3.8 Flash