Currently listed through these providers:
Model details
Gemini 2.5 Flash-Lite
Gemini 2.5 Flash-Lite is positioned as a lean member of the Gemini 2.5 family, aimed at cost-sensitive, high-scale workloads where latency and price matter more than top-tier reasoning depth. According to Google Cloud's announcement, the model is well suited for classification, translation, intelligent routing, and other production patterns that need to process large volumes efficiently on Vertex AI. Its design intent is to give enterprise builders a predictable, scalable option for routine language and multimodal tasks without invoking the heavier Flash or Pro tiers.
As part of the broader Gemini 2.5 release wave, Flash-Lite followed the general-availability milestones for Flash and Pro, giving organizations a tiered lineup that pairs reasoning strength with operational efficiency. Google described the rollout as focused on stability and reliability for production deployments, with Flash-Lite specifically highlighted for developers and enterprise builders who need to scale AI features economically. This positions the model as a practical building block for workflows where high throughput and low cost per call are the primary design constraints.
Quick Info
Powered by- Provider
- CrossModel
- Model key
- gemini/gemini-2.5-flash-lite
- Release date
- Jun 17, 2025
- Last updated
- Jun 17, 2025
- Knowledge cutoff
- 2025-01
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.10
- Output token cost
- $0.40
Limits
- Output tokens
- 65,536 tokens
- Context window
- 1,048,576 tokens
Transparent token rates
Compare Gemini 2.5 Flash-Lite pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Gemini 2.5 Flash-Lite
No articles yet. Fetch the latest news to show it here.
Videos about Gemini 2.5 Flash-Lite
More models around Gemini 2.5 Flash-Lite
This exact model name is also listed by 15 other providers.
