Currently listed through these providers:
Model details
Gemini 3.5 Flash Lite
Gemini 3.5 Flash-Lite joins the Gemini 3 series as a cost-efficient and fast addition, designed for high-volume, latency-sensitive workloads such as translation and classification. According to the official model card, it is a natively multimodal reasoning model that also supports agentic workflows, making it well suited for teams that need quick, structured outputs from diverse inputs without the overhead of larger flagship variants. Its positioning within the Flash-Lite family signals a focus on throughput and affordability rather than maximum reasoning depth.
The model is published with a dedicated model card on Google DeepMind's site, dated July 2026, alongside a Google Cloud Enterprise documentation entry, confirming its availability for production deployment through enterprise channels. The Flash-Lite lineage has historically emphasized efficient inference for classification and routing use cases, and this release continues that trajectory while bringing natively multimodal inputs and reasoning capabilities to cost-sensitive pipelines. For practitioners, it represents a practical choice when handling large volumes of requests where response speed and operating cost matter more than the deepest analytical reasoning.
Quick Info
Powered by- Provider
- NanoGPT
- Model key
- google/gemini-3.5-flash-lite
- Release date
- Jul 21, 2026
- Last updated
- Jul 21, 2026
- Knowledge cutoff
- 2026-03
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.30
- Output token cost
- $2.50
Limits
- Input tokens
- 1,048,576 tokens
- Output tokens
- 65,536 tokens
- Context window
- 1,048,576 tokens
Latest news about Gemini 3.5 Flash Lite
Videos about Gemini 3.5 Flash Lite
Recent tweets and retweets from NanoGPT
More models around Gemini 3.5 Flash Lite
This exact model name is also listed by 21 other providers.