Sulat.com
AI models
Google logo

Model details

Gemini 2.5 Flash-Lite

Gemini 2.5 Flash Lite is designed as the speed-optimized member of Google's Gemini 2.5 family, engineered for applications where response time and operational cost matter more than raw benchmark dominance. The model streams at approximately 393 tokens per second with a time-to-first-token around 0.29 seconds, making it one of the fastest production-ready options for high-volume, user-facing workloads. Its million-token context window lets developers feed entire documents, codebases, or books without chunking, while multimodal inputs handle text, images, audio, and video under a unified API. By default, the model disables multi-pass "thinking" to prioritize latency, though developers can selectively enable reasoning budgets when deeper analysis outweighs the speed penalty.

The model extends Google's Flash lineage, being roughly 1.5 times faster than its Gemini 2.0 predecessor while delivering improved performance across standard benchmarks. Its "intelligence per dollar" philosophy shaped the training approach—balancing quality with affordability for large-scale deployment scenarios like classification pipelines, data processing jobs, and high-traffic chatbots where premium models would be economically impractical. The optional reasoning toggle allows the same endpoint to handle both rapid Q&A and deliberate problem-solving tasks without model swapping, effectively blending two operational modes into one service. For teams building automated workflows or consumer applications where latency directly impacts user experience, Flash Lite offers a pragmatic path to production-grade AI without the cost ceiling of larger siblings.

Googlegemini-2.5-flash-litegemini-flash-lite

Quick Info

Powered by
Provider
Google
Model key
gemini-2.5-flash-lite
Release date
Jun 17, 2025
Last updated
Jun 17, 2025
Knowledge cutoff
2025-01
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.10
Output token cost
$0.40

Limits

Output tokens
65,536 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare gemini-flash-lite pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Gemini 2.5 Flash-Lite

Google

CoverageComparison

A data-driven comparison of Qwen3.5-Flash and Gemini 2.5 Flash-Lite - two models at the exact same $0.10/$0.40 per million token price point with 1M context windows but very different performance profiles.

Google

CoveragePreview

Google is pulling Gemini 2.5 Flash-Lite Preview from AI Studio on March 31. The replacement, Gemini 3.1 Flash-Lite Preview, costs significantly more per token.

Google

CoverageComparison

As of March 20, 2026, Gemini 2.5 Flash-Lite is still the better default if your main goal is the lowest stable token cost, while Gemini 3.1 Flash-Lite is the stronger successor lane if you can justify a much higher price for better quality and an eventual migration path. This guide explains when to stay, when to switch

Google

Coverage

The new model aims to address a significant challenge enterprise developers face by providing levels of thinking to better match the task at hand.

Videos about Gemini 2.5 Flash-Lite

More models around Gemini 2.5 Flash-Lite