Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
LLM Gateway logo

Model details

Gemini 3.8 Flash (Google Vertex AI)

Gemini 3.8 Flash sits inside Google's Gemini family of multimodal models and is positioned as a fast, lightweight option aimed at coding and agentic workloads. The LLM Gateway listing describes it as providing fast multimodal reasoning, and the entry is marked STABLE on the gateway's Vertex AI route. An August 28, 2026 Shattered.io report, citing Business Insider, states that Google staff have been running an internal preview referred to as Gemini 3.8 Flash Preview through Google's internal coding platform Jetski, just 14 days after Gemini 3.7 Flash reached general availability. That timeline helps explain the rapid point-release cadence inside the Flash line, where each iteration is intended to refine reasoning quality and developer ergonomics rather than introduce a new architecture.

In practical terms, Gemini 3.8 Flash is shaped for interactive developer scenarios: short-latency inference, multimodal input handling, and tool-using agent flows where quick iteration matters more than maximum depth. The Shattered.io coverage emphasizes its use inside Jetski, suggesting Google itself is leaning on the model to accelerate internal coding assistants. For external teams, that same profile translates well to chat assistants, IDE plugins, structured data extraction, and lightweight retrieval-augmented pipelines that need a responsive generalist. Buyers comparing it against larger Gemini tiers should expect a model tuned for speed and breadth, with the trade-off in raw reasoning depth that typically accompanies a Flash-class design.

LLM Gatewaygoogle-vertex/gemini-3.8-flashgemini

Quick Info

Powered by
Provider
LLM Gateway
Model key
google-vertex/gemini-3.8-flash
Release date
Sep 2, 2026
Last updated
Sep 2, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.75
Output token cost
$3.75

Limits

Output tokens
65,536 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare Gemini 3.8 Flash (Google Vertex AI) pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Gemini 3.8 Flash (Google Vertex AI)

LLM Gateway

Coverage

Tech Insider's coverage of the September 2, 2026 joint launch confirms that Google DeepMind shipped both Gemini 3.8 Flash and a specialized Gemini 3.8 Flash Cyber variant, with Google DeepMind's own X post stating "3.8 Flash: our most intelligent model yet with significant gains from 3.7 Flash across software engineeri Gemini 3.8 Flash reached general availability the same day through the Gemini API, Google AI Studio, Antigravity, Android Studio, and Gemini Enterprise, while the Cyber variant is gated behind a new Fairwind access program and will not be put on general release, mirroring the restricted rollout used for the earlier Gem

LLM Gateway

CoverageBenchmark

Armes.ai's deep dive on the September 2, 2026 release characterizes Gemini 3.8 Flash (codenamed "Skimaki") as a multimodal workhorse built on a sparse Mixture-of-Experts architecture with a native 1-million-token context window and up to 65,536 output tokens. The article cites Terminal-Bench 2.1 at 90.8% (versus 81.6% The page also discusses benchmark gains in execution-heavy domains such as terminal automation, real-world bug fixing, financial analysis, and scientific literature reasoning, and reports that internal Google developer telemetry on the Jetski coding platform showed engineers preferring 3.8 Flash over Claude Opus for da

LLM Gateway

CoverageRelease Notes

Google announced Gemini 3.8 Flash on September 2, 2026, its third Flash model release in roughly six weeks. The launch comes in two variations: the standard Gemini 3.8 Flash, described by Google as a "workhorse" model suited to agentic tasks and software development, and Gemini 3.8 Flash Cyber, tuned on the same founda API pricing for Gemini 3.8 Flash is set at an introductory rate of $0.75 per million input tokens and $3.75 per million output tokens, available through the end of 2026, with the standard rate rising to $1.50 input and $7.50 output per million tokens afterward. This matches the introductory pricing Google used for Gemi

LLM Gateway

CoverageBenchmark

ExplainX's post-launch summary, updated September 3, 2026, details Google-confirmed information for Gemini 3.8 Flash and its cybersecurity-focused sibling Gemini 3.8 Flash Cyber, referencing a joint launch post from Tulsee Doshi (Senior Director of Product Management) and Raluca Ada Popa (Gemini Security Lead, Google D The article also reports that Gemini 3.8 Flash Cyber is gated behind a new Fairwind access program and is not on general release, with the page noting CyberGym frontier-level scores and a CWE-Bench pass@1 of 47.2% versus 47.8% for the leading frontier model. This is useful corroboration of vendor-reported benchmark num

LLM Gateway

CoverageRelease Notes

Benchr's recent releases reference lists Google releasing the stable gemini-3.8-flash API model on September 2, 2026. The model has a 1,048,576-token input limit, a 65,536-token maximum output, and supports low, medium, or high thinking levels. Google lists introductory rates of $0.75 input and $3.75 output per million Benchr's September 2026 entry also covers GPT-6 Astra (September 3), Claude Fable 5.1 and Claude Mythos 5.1 (September 1), and an August 2026 GPT-5.6 Sol price cut. It notes API model records as the primary source for the Gemini 3.8 Flash spec details and links back to Google's launch materials and pricing documentatio

LLM Gateway

CoverageRelease Notes

GitHub announced on September 3, 2026 that Gemini 3.8 Flash is now available in GitHub Copilot. According to GitHub's early testing, the model performed strongly on complex terminal-based coding tasks and demonstrated rigorous validation with persistent recovery from actionable failures. The model is billed at introduc GitHub Copilot Enterprise and Business administrators control access to Gemini 3.8 Flash through model policy settings in Copilot. Under default model enablement, new models are turned on automatically unless a global default has been disabled or this specific model is explicitly turned off. GitHub directs users to its

LLM Gateway

CoverageBenchmark

Hokai's independent model hub page describes Gemini 3.8 Flash as Google DeepMind's September 2, 2026 Flash-tier release, further trained from Gemini 3.7 Flash on a sparse Mixture-of-Experts transformer with a 1M-token input window and 64K-token output ceiling. Independent testing by DataCamp shows gains over 3.7 Flash The hub explicitly lists the model's distribution channels as API, Google Vertex AI, and Google AI Studio, and notes practical tradeoffs such as a slow time-to-first-token averaging close to 13 seconds that make the model unsuitable for live chat or voice products. Exact parameter counts are not disclosed, consistent w

LLM Gateway

CoverageRelease Notes

Google's official Gemini API release notes for September 2, 2026 confirm that gemini-3.8-flash reached general availability, described by Google as "our most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows." The entry points developers to th Because this is the first-party Google AI for Developers documentation, the page directly names the exact model identifier gemini-3.8-flash and its GA status as of September 2, 2026. The excerpt does not separately call out Google Vertex AI as a distribution channel in this changelog entry, but it establishes the model

Videos about Gemini 3.8 Flash (Google Vertex AI)

More models around Gemini 3.8 Flash (Google Vertex AI)