Sulat.com
AI models
Merge Gateway logo

Model details

Gemini 3.1 Flash-Lite

Gemini 3.1 Flash Lite is Google's high-efficiency multimodal offering in the Flash family, presented on Google DeepMind's product page as a scalable thinking model aimed at high-volume workloads where cost and latency matter more than peak reasoning depth. It sits below Gemini 3 Flash in both capability and price, and OpenRouter explicitly notes that it is priced at half the cost of Gemini 3 Flash, reinforcing its positioning as a budget-tier default for production traffic. The model accepts a broad mix of inputs — text, images, video, audio, and PDF documents — while returning text, which makes it a natural fit for lightweight agentic loops, simple structured data extraction, UI generation, translation, and routing or classification layers in front of larger models.

A defining feature of Gemini 3.1 Flash Lite is its configurable thinking-level control, exposing minimal, low, medium, and high reasoning settings so developers can dial the trade-off between response speed, compute spend, and answer quality on a per-request basis. Combined with its million-token context window and generally available status, the model is designed for responsive, API-bound applications where throughput and predictability dominate the design constraints. In practice it is best matched to workloads such as bulk document and media summarization, conversation triage, code scaffolding, content transformation pipelines, and high-QPS assistants that need reliable multimodal understanding without paying flagship-model prices.

Merge Gatewaygoogle/gemini-3.1-flash-litegemini-flash-lite

Quick Info

Powered by
Provider
Merge Gateway
Model key
google/gemini-3.1-flash-lite
Release date
May 7, 2026
Last updated
May 7, 2026
Knowledge cutoff
2025-01
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.25
Output token cost
$1.50

Limits

Output tokens
65,536 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare Gemini 3.1 Flash-Lite pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Gemini 3.1 Flash-Lite

Videos about Gemini 3.1 Flash-Lite

More models around Gemini 3.1 Flash-Lite