Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
CrossModel logo

Model details

Gemini 3.5 Flash Lite

Gemini 3.5 Flash-Lite is positioned by Google DeepMind as its fastest and most cost-effective 3.5-class model, purpose-built for low-latency and high-throughput agentic tasks. It is generally available through the Gemini API alongside Gemini 3.6 Flash, with Google describing it as ready for production deployment. The model delivers 350 output tokens per second according to the Artificial Analysis Index, making it well suited to high-volume workloads such as extraction, search, translation, classification, and subagent orchestration. Its design emphasizes speed and economy over maximum reasoning depth, fitting naturally into pipelines where many lightweight calls need to happen quickly rather than a single deep inference pass.

In practical terms, Gemini 3.5 Flash-Lite is aimed at teams running large-scale pipelines that need responsive, affordable inference: coding assistance, UI generation, translation, and other repeatable agentic tasks. It supports a one-million-token context window with a 64k maximum output, native multimodal input, thinking controls, and built-in tools, giving it flexibility across varied task shapes while keeping per-call costs low. The combination of high throughput, generous context, and production-ready status makes it a natural choice for serving as the lightweight tier in a multi-model setup, handling the bulk of routine requests so heavier models can be reserved for the hardest problems.

CrossModelgemini/gemini-3.5-flash-litegemini-flash-lite

Quick Info

Powered by
Provider
CrossModel
Model key
gemini/gemini-3.5-flash-lite
Release date
Jul 21, 2026
Last updated
Jul 21, 2026
Knowledge cutoff
2026-03
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.30
Output token cost
$2.50

Limits

Output tokens
65,536 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare Gemini 3.5 Flash Lite pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Gemini 3.5 Flash Lite

CrossModel

CoverageRelease Notes

Opper AI's Google model release tracker explicitly lists Gemini 3.5 Flash Lite with a 21 Jul 2026 release date, a 1M-token context window, $0.30 input / $2.50 output per million tokens pricing, and an intelligence index of 37. This corroborates the release timing and pricing reported by BenchLM for the exact model vari Because Opper compiles Google release metadata rather than publishing first-party technical details, this entry functions as a secondary corroboration of the model's existence, launch window, and price points rather than an announcement of new capabilities. No benchmark detail, paper, or developer documentation is incl

CrossModel

Coverage

Google began rolling out Gemini 3.5 Flash-Lite inside Google Search following the official July 21, 2026 launch announcement, with this third-party news article dated July 27, 2026. Among the three models announced that day (Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber), only 3.5 Flash-Lite was p Vendor-reported benchmark numbers from Google's announcement are quoted: SWE-Bench Pro at 54.2% versus 3 Flash's 49.6%, and OSWorld-Verified at 74.0% versus 3 Flash's 65.1%. The article explicitly notes these figures come from Google's own testing rather than independent reproduction. Companion details for the sibling

CrossModel

Coverage

Google shipped Gemini 3.5 Flash-Lite on July 21, 2026 alongside Gemini 3.6 Flash and the security-focused Gemini 3.5 Flash Cyber variant, according to third-party coverage of Google's official announcement. Gemini 3.5 Flash-Lite is positioned as Google's fastest and cheapest tier in the Gemini family, targeted at high- The article frames the launch as Google's response to mid-tier refreshes from OpenAI, Anthropic, and xAI, with Gemini 3.6 Flash replacing Gemini 3.5 Flash as the workhorse at $1.50 per million input tokens and $7.50 per million output tokens. Gemini 3.5 Pro did not ship in this release and remains in partner testing pe

CrossModel

CoverageBenchmark

An educational blog post dated July 21, 2026 summarizes Google's GA launch of Gemini 3.6 Flash and Gemini 3.5 Flash-Lite, alongside the specialized Gemini 3.5 Flash Cyber vulnerability-discovery system. Gemini 3.5 Flash-Lite is described as a high-throughput, low-latency model targeted at extraction, search, translatio The post attributes its spec sheet to Google's official three-model announcement, the Gemini API latest-model guide, and the respective Gemini 3.6 Flash and Gemini 3.5 Flash-Lite model cards, with a source-check date of July 21, 2026. Sibling Gemini 3.6 Flash is positioned as the general agentic, coding, knowledge, and

CrossModel

CoverageBenchmark

An independent model-tracking page for "Gemini 3.5 Flash-Lite" (dated July 21, 2026) provides an aggregator view of the model's positioning, ranking it 127 on the LLM Stats composite score and reporting a blended price of $0.40 per million tokens against a score of 29.8 on the site's leaderboard. The page tabulates ben Quality-tracker data shows stability around the baseline with a 34-vote 7-day sample, and the page benchmarks Gemini 3.5 Flash-Lite against comparators including Gemma 4 E4B, GPT OSS 120B, DeepSeek-V4-Flash-0731, DeepSeek-V4.1-Flash, and GPT-6 Astra on a log-scale cost-versus-score plot. The site is an independent aggr

CrossModel

CoverageBenchmark

The BenchLM model profile documents Gemini 3.5 Flash Lite as a Google proprietary model released on July 21, 2026, with a 1-million-token context window, a listed API price of $0.30 per million input tokens and $2.50 per million output tokens, and a cached-input price of $0.03 per million tokens. The profile also repor On the same profile, Gemini 3.5 Flash Lite is assigned a capability score of 59 out of 100, ranking it 65th of 231 tracked models, with measured throughput of 378 tokens per second and a first-token latency of 7.35 seconds. The model's strongest ranked category is Knowledge at 68, while Coding is its lowest eligible ca

Videos about Gemini 3.5 Flash Lite

More models around Gemini 3.5 Flash Lite