Currently listed through these providers:
Model details
gemini-3.1-flash-lite
Gemini 3.1 Flash-Lite is positioned as the generally available efficiency tier of the Gemini 3.1 generation, building on the lineage established by 2.5 Flash Lite with targeted improvements in reasoning, multimodal understanding, agentic tool use, and long-context behavior. It accepts text, image, video, audio, and PDF inputs and is described as a low-latency, cost-effective multimodal model aimed at high-volume, lightweight tasks such as simple data extraction and frequent API calls where speed and price outweigh the need for deeper deliberation.
For developers, the model exposes four configurable thinking levels so teams can trade depth of reasoning against response latency on a per-request basis, and it operates within a roughly one-million-token context window that supports long-document and multi-turn agentic workflows. Capability coverage surfaced for the model includes reasoning, tool use, implicit caching, file input handling, vision and image understanding, and web search, making it a practical fit for retrieval-augmented assistants and pipeline-style automation. Together, these traits make Gemini 3.1 Flash-Lite best suited to production scenarios where consistent throughput and predictable cost matter more than frontier reasoning quality.
Quick Info
Powered by- Provider
- SAP AI Core
- Model key
- gemini-3.1-flash-lite
- Release date
- May 7, 2026
- Last updated
- May 7, 2026
- Knowledge cutoff
- 2025-01
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.25
- Output token cost
- $1.50
Limits
- Output tokens
- 65,536 tokens
- Context window
- 1,048,576 tokens
Latest news about gemini-3.1-flash-lite
Videos about gemini-3.1-flash-lite
Recent tweets and retweets from SAP AI Core
More models around gemini-3.1-flash-lite
This exact model name is also listed by 21 other providers.