Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Vertex logo

Model details

Gemini 3.5 Flash Lite

Gemini 3.5 Flash-Lite is documented on Google's first-party Gemini API site and promoted through a dedicated Google DeepMind product page, positioning it inside the wider Gemini family of large language models. Google DeepMind describes the variant as the fastest and most cost-effective option in the 3.5 class, explicitly aimed at low-latency, high-throughput agentic use cases such as coding assistance, UI generation, and translation where rapid responses and high request volume matter more than maximum reasoning depth.

Google DeepMind's page cites a measured output speed of around 350 tokens per second for Gemini 3.5 Flash-Lite, attributed to the Artificial Analysis Index, which signals that the model is engineered for token-efficient streaming rather than frontier reasoning. The product page is organized into capabilities, hands-on, showcase, performance, and model-information sections, and links directly into AI Studio with the model identifier gemini-3.5-flash-lite, making it straightforward for developers to prototype agent pipelines that need quick, inexpensive responses at scale.

Vertexgemini-3.5-flash-litegemini-flash-lite

Quick Info

Powered by
Provider
Vertex
Model key
gemini-3.5-flash-lite
Release date
Jul 21, 2026
Last updated
Jul 21, 2026
Knowledge cutoff
2026-03
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.30
Output token cost
$2.50

Limits

Output tokens
65,536 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare Gemini 3.5 Flash Lite pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Gemini 3.5 Flash Lite

Vertex

Official sourceDocumentation

Google Cloud's official documentation page describes Gemini 3.5 Flash-Lite as the latest model in the cost-effective Flash-Lite line, optimized for simple coding tasks, precise document understanding, and lightweight agentic workflows that need fast inference at minimal cost. It is positioned as a suitable replacement The page also enumerates breaking changes versus prior Gemini models: custom values for temperature, top-K, and top-P are ignored; custom values for frequency and presence penalty throw errors; and requests where the last input turn has a Model role (such as "type": "model output" in the Interactions API or "role": "mo

302.AI

CoverageRelease Notes

Google's official Gemini API release notes show a September 1, 2026 update adding agentic video understanding for Gemini 3.5 Flash-Lite (alongside 3.7 and 3.6 Flash) across the Interactions and GenerateContent APIs. The model dynamically navigates video timelines, requesting transcripts, frames, or audio tracks on dema The changelog entry names the exact variant "Gemini 3.5 Flash-Lite" as a recipient of the new video-understanding capability, satisfying strict version discipline. It also flags newer sibling models (3.6 Flash, 3.7 Flash, 3.8 Flash) for context, but attributes them to their own model IDs rather than conflating them wit

SAP AI Core

Coverage

Google began rolling out Gemini 3.5 Flash-Lite inside Google Search as of late July 2026, following its official announcement on July 21, 2026 alongside Gemini 3.6 Flash and Gemini 3.5 Flash Cyber, according to almcorp.com. Gemini 3.5 Flash-Lite is described as the fastest and cheapest model in the Gemini family, targe On benchmarks, Google claims Gemini 3.5 Flash-Lite scored 54.2% on SWE-Bench Pro versus 3 Flash's 49.6%, and 74.0% on OSWorld-Verified versus 3 Flash's 65.1% (figures from Google's own announcement). Among the three models released July 21, only Gemini 3.5 Flash-Lite is being rolled out to Google Search, while Gemini 3

Opper

Coverage

Tech Insider reported on July 23, 2026 that Google shipped three new Gemini models on July 21, 2026: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and a narrow security-focused variant called Gemini 3.5 Flash Cyber. The article confirms Gemini 3.5 Flash-Lite's distribution across the Gemini API through Google AI Studio and The piece frames the release as Google's answer for the price-sensitive, high-volume segment of the API market, where cost per token decides which model wins contracts for coding agents, support bots, and search features. Gemini 3.5 Pro did not ship in this release and remains in partner testing with general availabili

SAP AI Core

Coverage

Technology.org reported on July 23, 2026 that Google released a trio of cheaper Gemini models on Tuesday covering coding, cost efficiency, and cybersecurity: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and a new cybersecurity-focused variant called Gemini 3.5 Flash Cyber. The three models target workloads that do not need The same Technology.org piece notes that Gemini 3.5 Pro was originally slated for a June launch, announced by CEO Sundar Pichai at the company's I/O developer conference in May, but did not ship in this release. The delay is attributed to the model falling short of internal goals, particularly on coding, which has beco

Abacus

CoverageBenchmark

Coursiv's blog documents Google's July 21, 2026 launch of Gemini 3.6 Flash and Gemini 3.5 Flash-Lite as generally available production models, alongside the specialized Gemini 3.5 Flash Cyber system for vulnerability discovery and patching inside CodeMender. Gemini 3.5 Flash-Lite is positioned for high-volume, low-late The article explicitly cites verified Google sources checked July 21, 2026: the official three-model announcement, the Gemini API latest-model guide for production model IDs and code examples, the Gemini 3.6 Flash model card, and the Gemini 3.5 Flash-Lite model card for context, output, cutoff, intended uses, and safet

Impossibl

Coverage

Ars Technica's July 21, 2026 coverage of the Google announcement explicitly names Gemini 3.5 Flash Lite and confirms its positioning as Google's most efficient modern AI model, citing the same 350 tokens-per-second figure. The article places this release in the context of the Gemini 3.5 Flash deprecation and the simult Independent reporting treats 3.5 Flash Lite as ideal for scaling agentic systems without breaking the bank. Benchmarks cited in the piece indicate the new Flash Lite is roughly comparable to or slightly below the outgoing 3.5 Flash on certain tests, with the story's center of gravity on Gemini 3.6 Flash. Only the Flash

Videos about Gemini 3.5 Flash Lite

More models around Gemini 3.5 Flash Lite