Google is pulling Gemini 2.5 Flash-Lite Preview from AI Studio on March 31. The replacement, Gemini 3.1 Flash-Lite Preview, costs significantly more per token.
Model details
Gemini 3.1 Flash Lite Preview
Gemini 3.1 Flash Lite Preview represents Google's focused effort to build a high-efficiency model for high-volume production workloads. Rather than chasing reasoning-heavy capabilities, this model prioritizes consistent, repeatable performance across translation, tagging, and content moderation tasks that developers run millions of times daily. The architecture is optimized to avoid the massive compute overhead of larger reasoning models, making it practical to deploy at scale through both Gemini API and Vertex AI. This design philosophy positions the model as a workhorse for applications where reliability and throughput matter more than extensive chain-of-thought reasoning.
The model builds on Google's Flash lineage, delivering measurable improvements over its 2.5 Flash Lite predecessor across audio input and automatic speech recognition, RAG snippet ranking, translation quality, data extraction accuracy, and code completion. Notably, it approaches the performance of the larger Gemini 2.5 Flash across key benchmarks while maintaining a leaner profile. Users can also dial in thinking effort across minimal, low, medium, and high levels, enabling fine-grained control over the cost-performance tradeoff for different task types. At half the cost of Gemini 3 Flash, it fills a strategic niche for developers who need reliable everyday AI capabilities without the premium pricing of flagship models.
Quick Info
Powered by- Provider
- Model key
- gemini-3.1-flash-lite-preview
- Release date
- Mar 3, 2026
- Last updated
- Mar 3, 2026
- Knowledge cutoff
- 2025-01
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.25
- Output token cost
- $1.50
Limits
- Output tokens
- 65,536 tokens
- Context window
- 1,048,576 tokens
Transparent token rates
Compare gemini-flash-lite pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Gemini 3.1 Flash Lite Preview
Gemini 3.1 Flash Lite Preview is Google's high-efficiency model optimized for high-volume use cases. $0.25 per million input tokens, $1.50 per million output tokens. 1,048,576 token context window, maximum output of 65,536 tokens. Higher uptime with 2 providers.
Videos about Gemini 3.1 Flash Lite Preview
More models around Gemini 3.1 Flash Lite Preview
This exact model name is also listed by 9 other providers.