Currently listed through these providers:
Model details
Gemini 3.1 Flash-Lite
Gemini 3.1 Flash Lite is Google's high-efficiency multimodal offering in the Flash family, presented on Google DeepMind's product page as a scalable thinking model aimed at high-volume workloads where cost and latency matter more than peak reasoning depth. It sits below Gemini 3 Flash in both capability and price, and OpenRouter explicitly notes that it is priced at half the cost of Gemini 3 Flash, reinforcing its positioning as a budget-tier default for production traffic. The model accepts a broad mix of inputs — text, images, video, audio, and PDF documents — while returning text, which makes it a natural fit for lightweight agentic loops, simple structured data extraction, UI generation, translation, and routing or classification layers in front of larger models.
A defining feature of Gemini 3.1 Flash Lite is its configurable thinking-level control, exposing minimal, low, medium, and high reasoning settings so developers can dial the trade-off between response speed, compute spend, and answer quality on a per-request basis. Combined with its million-token context window and generally available status, the model is designed for responsive, API-bound applications where throughput and predictability dominate the design constraints. In practice it is best matched to workloads such as bulk document and media summarization, conversation triage, code scaffolding, content transformation pipelines, and high-QPS assistants that need reliable multimodal understanding without paying flagship-model prices.
Quick Info
Powered by- Provider
- Merge Gateway
- Model key
- google/gemini-3.1-flash-lite
- Release date
- May 7, 2026
- Last updated
- May 7, 2026
- Knowledge cutoff
- 2025-01
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.25
- Output token cost
- $1.50
Limits
- Output tokens
- 65,536 tokens
- Context window
- 1,048,576 tokens
Transparent token rates
Compare Gemini 3.1 Flash-Lite pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Gemini 3.1 Flash-Lite
Videos about Gemini 3.1 Flash-Lite
More models around Gemini 3.1 Flash-Lite
This exact model name is also listed by 21 other providers.