Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Vercel AI Gateway logo

Model details

Gemini 2.5 Flash Lite

Gemini 2.5 Flash Lite is a lightweight thinking model in Google's Gemini 2.5 family, built to deliver fast, cost-efficient responses without sacrificing quality. Unlike standard language models that generate responses in a single pass, thinking models internally reason through their thoughts before answering, which can improve accuracy on complex tasks. Flash Lite is designed to be lean and fast by default, keeping reasoning disabled so applications can prioritize speed and low latency. Developers who encounter tougher problems can opt into multi-pass reasoning through the Reasoning API, giving them a knob to trade off speed for deeper analysis when needed.

This model builds on the lineage of earlier Flash models while offering measurably better performance across common benchmarks and faster token generation. The thinking budget control, shared across the Gemini 2.5 family, means developers can tune how much computational reasoning the model applies on a per-query basis. This flexibility makes Flash Lite practical for high-throughput production environments where most requests are straightforward, while still allowing intelligent fallback to deeper reasoning for edge cases. It's positioned as an accessible entry point into the Gemini 2.5 thinking ecosystem, making advanced reasoning capabilities practical for applications that can't afford the latency or cost of heavier models.

Vercel AI Gatewaygoogle/gemini-2.5-flash-litegemini-flash-lite

Quick Info

Powered by
Provider
Vercel AI Gateway
Model key
google/gemini-2.5-flash-lite
Release date
Jun 17, 2025
Last updated
Jun 17, 2025
Knowledge cutoff
2025-01
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.10
Output token cost
$0.40

Limits

Output tokens
65,535 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare Gemini 2.5 Flash Lite pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Gemini 2.5 Flash Lite

Vercel AI Gateway

Coverage

Google has been on a roll lately with its Gemini lineup of large language models, and now the company is expanding the 2.5 family with a new addition., Google has been on a roll lately with its Gemini lineup of large language models, and now the company is expanding the 2.5 family with a new addition.

Vercel AI Gateway

CoverageBenchmark

Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. $0.10 per million input tokens, $0.40 per million output tokens. 1,048,576 token context window, maximum output of 65,535 tokens. Higher uptime with 2 providers. Includes independent ben

Videos about Gemini 2.5 Flash Lite

More models around Gemini 2.5 Flash Lite