Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Venice AI logo

Model details

Gemini 3.5 Flash-Lite

Gemini 3.5 Flash-Lite is positioned within Google's 3.x Flash family as the lightweight counterpart that pairs low latency with cost-effective throughput for production workloads. According to the official Gemini API changelog, it reached general availability on July 21, 2026 alongside Gemini 3.6 Flash, as part of a stable, production-ready release of the latest Flash-tier lineup. The model is explicitly framed as "a low-latency, highly cost-effective subagent option designed for high-volume automation," which signals that Google optimized the Flash-Lite branch for scaling many parallel, focused tasks rather than for single-shot deep reasoning.

In practice, the variant is described as a high-efficiency model with upgraded agentic capabilities, well suited to subagents that handle discrete responsibilities inside larger multi-agent systems. Vertex AI listings confirm a one-the cataloged API limit and a roughly 66,000-token maximum output, giving subagents enough room to consume long tool traces, retrieval snippets, or code context while still returning substantial structured responses. The capability surface—spanning vision, reasoning, tool calling, caching, web search, and JSON-schema structured output—makes it flexible enough to plug into orchestrator-driven pipelines where one Flash-Lite instance might classify, route, or summarize before handing richer generation off to a larger model in the same agent graph.

Venice AIgemini-3-5-flash-litegemini-flash-lite

Quick Info

Powered by
Provider
Venice AI
Model key
gemini-3-5-flash-lite
Release date
Jul 9, 2026
Last updated
Jul 21, 2026
Knowledge cutoff
2026-03
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.375
Output token cost
$3.125

Limits

Output tokens
65,536 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare Gemini 3.5 Flash-Lite pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Gemini 3.5 Flash-Lite

Venice AI

Coverage

A July 27, 2026 report covers Google's rollout of Gemini 3.5 Flash-Lite inside Google Search, the only distribution channel of the three new models that gets Search access. It notes that Flash-Lite was officially announced on July 21, 2026 alongside Gemini 3.6 Flash and Gemini 3.5 Flash Cyber, and describes Flash-Lite The same piece reports Google's claim that Gemini 3.5 Flash-Lite scores 54.2% on SWE-Bench Pro versus 49.6% for Gemini 3 Flash, and 74.0% on OSWorld-Verified versus 65.1% for Gemini 3 Flash, while flagging that these are Google's own internal figures rather than independent confirmation. It also restates Gemini 3.6 Fla

Venice AI

CoverageBenchmark

Coursiv's spec-table write-up confirms Gemini 3.5 Flash-Lite launched as generally available on July 21, 2026, priced at $0.30 per million input tokens and $2.50 per million output tokens, targeting high-throughput extraction, search, translation, classification, and subagent workloads. The page lists a 1-million-token The same article pairs Flash-Lite with Gemini 3.6 Flash, the workhorse replacement for 3.5 Flash that uses 17% fewer output tokens and lowers output pricing from $9 to $7.50 per million tokens, and with Gemini 3.5 Flash Cyber, a limited-pilot vulnerability model without self-serve API pricing. It links to Google's offi

Venice AI

Coverage

Ars Technica's July 21, 2026 report on Google's three-model announcement describes Gemini 3.5 Flash-Lite as the company's "most efficient modern AI," reporting a throughput of approximately 350 tokens per second and positioning it for scaling agentic systems. The article places Flash-Lite alongside Gemini 3.6 Flash (th For model-level news, this independent third-party reporting corroborates Google's framing of Flash-Lite as the speed- and cost-optimized sibling of 3.6 Flash, and it confirms the July 21, 2026 GA date that the Coursiv candidate also references. The candidate is weakest on standalone Flash-Lite technical detail because

Venice AI

Coverage

On July 21, 2026, Google released Gemini 3.5 Flash-Lite alongside Gemini 3.6 Flash and the security-tuned Gemini 3.5 Flash Cyber, positioning Flash-Lite as the cheapest, lowest-latency option in the Gemini family. The candidate describes it as a clear step up from the 3.1 Flash-Lite it replaces and notes it is aimed at The same report frames the Flash-tier push as happening while Gemini 3.5 Pro stays in limited partner testing and Google has begun what it calls its most ambitious pre-training run yet for Gemini 4. It places Flash-Lite at the speed-optimized end of the lineup, with Flash-Lite specifically called out as rolling into Se

Venice AI

Coverage

Google's July 21, 2026 announcement page details Gemini 3.5 Flash-Lite as the fastest, most cost-effective 3.5-class model in the drop, generally available the same day. It cites 350 output tokens per second from the Artificial Analysis Index and frames Flash-Lite around higher token efficiency, lower latency, and more The same article situates Flash-Lite next to Gemini 3.6 Flash, a new workhorse that uses 17% fewer output tokens than 3.5 Flash and adds built-in Computer Use tooling, and Gemini 3.5 Flash Cyber, a limited-access vulnerability-finding model. It also reiterates that Gemini 3.5 Pro is still in partner testing and that a

Venice AI

CoverageRelease Notes

Google's official Gemini API release notes (ai.google.dev/gemini-api/docs/changelog) document a September 1, 2026 update that extends agentic video understanding to Gemini 3.5 Flash-Lite alongside 3.7 Flash and 3.6 Flash, available across the Interactions and GenerateContent APIs. The feature lets the model dynamically For developers using the Gemini API (the same surface through which Venice AI exposes the gemini-3-5-flash-lite model key), the September 1 changelog entry confirms that Flash-Lite continues to receive first-party capability updates rather than being frozen at GA. Because the evidence comes directly from Google's API c

Videos about Gemini 3.5 Flash-Lite

More models around Gemini 3.5 Flash-Lite