On June 17, 2025, Google introduced Gemini 2.5 Flash-Lite in preview alongside the general availability of Gemini 2.5 Flash and 2.5 Pro. Flash-Lite is described as the most cost-efficient and fastest member of the 2.5 family, designed for high-volume, latency-sensitive workloads like translation and classification. The preview launched simultaneously in Google AI Studio and Vertex AI. Flash-Lite offers higher quality than 2.0 Flash-Lite across coding, math, science, reasoning, and multimodal benchmarks, with lower latency than both 2.0 Flash-Lite and 2.0 Flash on a broad sample of prompts. Capabilities include toggleable thinking budgets, tool integration with Google Search and code execution, multimodal input, and a 1 million-token context window. Google also deployed custom versions of Flash-Lite and Flash into Search.
Model details
Gemini 2.5 Flash-Lite
Gemini 2.5 Flash-Lite sits inside the Gemini 2.5 family as its quick, efficiency-oriented sibling, engineered by Google DeepMind for production scenarios where responsiveness and scale matter more than top-tier reasoning depth. The model is positioned for latency-sensitive applications such as large-scale chatbots, classification pipelines, and data-processing workloads that benefit from a lightweight footprint. Its multimodal input surface spans text, images, audio, and video through a single API, while outputs remain text-only, which keeps downstream integration simple for teams building conversational and retrieval-style systems on top of long context sources.
Practically, Gemini 2.5 Flash-Lite is reported by third-party evaluators to stream at roughly 392 tokens per second with about a 0.29-second time-to-first-token, placing it among the fastest production-grade models currently available. It supports a one-million-token context window, allowing whole books, long PDFs, or extensive codebases to be ingested without manual chunking, and offers an optional thinking-budget mode that lifts math accuracy on benchmarks like AIME while still improving code generation. Compared to the prior Gemini 2.0 Flash generation, it is described as about 1.5 times faster, making it a pragmatic choice when teams want Gemini-family multimodal capabilities and very large context at a lower per-token price point.
Quick Info
Powered by- Provider
- NEAR AI Cloud
- Model key
- google/gemini-2.5-flash-lite
- Release date
- Jun 17, 2025
- Last updated
- Jun 17, 2025
- Knowledge cutoff
- 2025-01
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.10
- Output token cost
- $0.40
Limits
- Output tokens
- 65,536 tokens
- Context window
- 1,048,576 tokens
Transparent token rates
Compare Gemini 2.5 Flash-Lite pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Gemini 2.5 Flash-Lite
Google announced the stable, generally available release of Gemini 2.5 Flash-Lite on July 22, 2025, positioning it as the fastest and lowest-cost model in the Gemini 2.5 family. Pricing is set at $0.10 per 1M input tokens and $0.40 per 1M output tokens, with audio input pricing reduced 40% from the preview. The model targets latency-sensitive, high-volume tasks such as translation and classification. The stable release includes a 1 million-token context window, controllable thinking budgets, and native tool support for Grounding with Google Search, Code Execution, and URL Context. Google highlights benchmarks showing all-around quality gains over 2.0 Flash-Lite in coding, math, science, reasoning, and multimodal understanding. Deployment examples include Satlyt cutting onboard diagnostic latency by 45% and reducing power consumption by 30%, and HeyGen using the model for avatar creation workflows.
DevPass (LLM Gateway)
Google announced the stable release of Gemini 2.5 Pro and Flash and introduced a new Gemini 2.5 Flash-Lite variant in preview as part of the broader Gemini 2.5 family rollout. Flash-Lite is positioned as the most cost-effective and fastest option in the Gemini 2.5 lineup, designed as a cost-effective upgrade over the p According to the article, Gemini 2.5 Flash-Lite delivers higher overall quality than 2.0 Flash-Lite in programming, mathematics, science, reasoning, and multimodal benchmarks, and is especially tuned for high-volume, latency-sensitive tasks such as translation and classification, where it achieves lower latency than 2.
Videos about Gemini 2.5 Flash-Lite
More models around Gemini 2.5 Flash-Lite
This exact model name is also listed by 15 other providers.
