Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
NEAR AI Cloud logo

Model details

Gemini 2.5 Flash-Lite

Gemini 2.5 Flash-Lite sits inside the Gemini 2.5 family as its quick, efficiency-oriented sibling, engineered by Google DeepMind for production scenarios where responsiveness and scale matter more than top-tier reasoning depth. The model is positioned for latency-sensitive applications such as large-scale chatbots, classification pipelines, and data-processing workloads that benefit from a lightweight footprint. Its multimodal input surface spans text, images, audio, and video through a single API, while outputs remain text-only, which keeps downstream integration simple for teams building conversational and retrieval-style systems on top of long context sources.

Practically, Gemini 2.5 Flash-Lite is reported by third-party evaluators to stream at roughly 392 tokens per second with about a 0.29-second time-to-first-token, placing it among the fastest production-grade models currently available. It supports a one-million-token context window, allowing whole books, long PDFs, or extensive codebases to be ingested without manual chunking, and offers an optional thinking-budget mode that lifts math accuracy on benchmarks like AIME while still improving code generation. Compared to the prior Gemini 2.0 Flash generation, it is described as about 1.5 times faster, making it a pragmatic choice when teams want Gemini-family multimodal capabilities and very large context at a lower per-token price point.

NEAR AI Cloudgoogle/gemini-2.5-flash-litegemini-flash-lite

Quick Info

Powered by
Provider
NEAR AI Cloud
Model key
google/gemini-2.5-flash-lite
Release date
Jun 17, 2025
Last updated
Jun 17, 2025
Knowledge cutoff
2025-01
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.10
Output token cost
$0.40

Limits

Output tokens
65,536 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare Gemini 2.5 Flash-Lite pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Gemini 2.5 Flash-Lite

Google

Official sourceOfficial

On June 17, 2025, Google introduced Gemini 2.5 Flash-Lite in preview alongside the general availability of Gemini 2.5 Flash and 2.5 Pro. Flash-Lite is described as the most cost-efficient and fastest member of the 2.5 family, designed for high-volume, latency-sensitive workloads like translation and classification. The preview launched simultaneously in Google AI Studio and Vertex AI. Flash-Lite offers higher quality than 2.0 Flash-Lite across coding, math, science, reasoning, and multimodal benchmarks, with lower latency than both 2.0 Flash-Lite and 2.0 Flash on a broad sample of prompts. Capabilities include toggleable thinking budgets, tool integration with Google Search and code execution, multimodal input, and a 1 million-token context window. Google also deployed custom versions of Flash-Lite and Flash into Search.

Google

Official sourceDocumentation

Google announced the stable, generally available release of Gemini 2.5 Flash-Lite on July 22, 2025, positioning it as the fastest and lowest-cost model in the Gemini 2.5 family. Pricing is set at $0.10 per 1M input tokens and $0.40 per 1M output tokens, with audio input pricing reduced 40% from the preview. The model targets latency-sensitive, high-volume tasks such as translation and classification. The stable release includes a 1 million-token context window, controllable thinking budgets, and native tool support for Grounding with Google Search, Code Execution, and URL Context. Google highlights benchmarks showing all-around quality gains over 2.0 Flash-Lite in coding, math, science, reasoning, and multimodal understanding. Deployment examples include Satlyt cutting onboard diagnostic latency by 45% and reducing power consumption by 30%, and HeyGen using the model for avatar creation workflows.

DevPass (LLM Gateway)

Coverage

Google announced the stable release of Gemini 2.5 Pro and Flash and introduced a new Gemini 2.5 Flash-Lite variant in preview as part of the broader Gemini 2.5 family rollout. Flash-Lite is positioned as the most cost-effective and fastest option in the Gemini 2.5 lineup, designed as a cost-effective upgrade over the p According to the article, Gemini 2.5 Flash-Lite delivers higher overall quality than 2.0 Flash-Lite in programming, mathematics, science, reasoning, and multimodal benchmarks, and is especially tuned for high-volume, latency-sensitive tasks such as translation and classification, where it achieves lower latency than 2.

Videos about Gemini 2.5 Flash-Lite

More models around Gemini 2.5 Flash-Lite