Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
AIHubMix logo

Model details

Gemini 2.5 Flash-Lite

Gemini 2.5 Flash-Lite sits inside the Gemini 2.5 family as its quick, efficiency-oriented sibling, engineered by Google DeepMind for production scenarios where responsiveness and scale matter more than top-tier reasoning depth. The model is positioned for latency-sensitive applications such as large-scale chatbots, classification pipelines, and data-processing workloads that benefit from a lightweight footprint. Its multimodal input surface spans text, images, audio, and video through a single API, while outputs remain text-only, which keeps downstream integration simple for teams building conversational and retrieval-style systems on top of long context sources.

Practically, Gemini 2.5 Flash-Lite is reported by third-party evaluators to stream at roughly 392 tokens per second with about a 0.29-second time-to-first-token, placing it among the fastest production-grade models currently available. It supports a one-million-token context window, allowing whole books, long PDFs, or extensive codebases to be ingested without manual chunking, and offers an optional thinking-budget mode that lifts math accuracy on benchmarks like AIME while still improving code generation. Compared to the prior Gemini 2.0 Flash generation, it is described as about 1.5 times faster, making it a pragmatic choice when teams want Gemini-family multimodal capabilities and very large context at a lower per-token price point.

AIHubMixgemini-2.5-flash-litegemini-flash-lite

Quick Info

Powered by
Provider
AIHubMix
Model key
gemini-2.5-flash-lite
Release date
Jun 17, 2025
Last updated
Jun 17, 2025
Knowledge cutoff
2025-01
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.10
Output token cost
$0.40

Limits

Output tokens
65,536 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare Gemini 2.5 Flash-Lite pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Gemini 2.5 Flash-Lite

DevPass (LLM Gateway)

Coverage

Google announced the stable release of Gemini 2.5 Pro and Flash and introduced a new Gemini 2.5 Flash-Lite variant in preview as part of the broader Gemini 2.5 family rollout. Flash-Lite is positioned as the most cost-effective and fastest option in the Gemini 2.5 lineup, designed as a cost-effective upgrade over the p According to the article, Gemini 2.5 Flash-Lite delivers higher overall quality than 2.0 Flash-Lite in programming, mathematics, science, reasoning, and multimodal benchmarks, and is especially tuned for high-volume, latency-sensitive tasks such as translation and classification, where it achieves lower latency than 2.

Videos about Gemini 2.5 Flash-Lite

More models around Gemini 2.5 Flash-Lite