Kilo Gateway
A data-driven comparison of Qwen3.5-Flash and Gemini 2.5 Flash-Lite - two models at the exact same $0.10/$0.40 per million token price point with 1M context windows but very different performance profiles.
Model details
Gemini 2.5 Flash-Lite sits inside the Gemini 2.5 family as its quick, efficiency-oriented sibling, engineered by Google DeepMind for production scenarios where responsiveness and scale matter more than top-tier reasoning depth. The model is positioned for latency-sensitive applications such as large-scale chatbots, classification pipelines, and data-processing workloads that benefit from a lightweight footprint. Its multimodal input surface spans text, images, audio, and video through a single API, while outputs remain text-only, which keeps downstream integration simple for teams building conversational and retrieval-style systems on top of long context sources.
Practically, Gemini 2.5 Flash-Lite is reported by third-party evaluators to stream at roughly 392 tokens per second with about a 0.29-second time-to-first-token, placing it among the fastest production-grade models currently available. It supports a one-million-token context window, allowing whole books, long PDFs, or extensive codebases to be ingested without manual chunking, and offers an optional thinking-budget mode that lifts math accuracy on benchmarks like AIME while still improving code generation. Compared to the prior Gemini 2.0 Flash generation, it is described as about 1.5 times faster, making it a pragmatic choice when teams want Gemini-family multimodal capabilities and very large context at a lower per-token price point.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Kilo Gateway
A data-driven comparison of Qwen3.5-Flash and Gemini 2.5 Flash-Lite - two models at the exact same $0.10/$0.40 per million token price point with 1M context windows but very different performance profiles.
Kilo Gateway
Google is pulling Gemini 2.5 Flash-Lite Preview from AI Studio on March 31. The replacement, Gemini 3.1 Flash-Lite Preview, costs significantly more per token.
Kilo Gateway
As of March 20, 2026, Gemini 2.5 Flash-Lite is still the better default if your main goal is the lowest stable token cost, while Gemini 3.1 Flash-Lite is the stronger successor lane if you can justify a much higher price for better quality and an eventual migration path. This guide explains when to stay, when to switch
DevPass (LLM Gateway)
Google announced the stable release of Gemini 2.5 Pro and Flash and introduced a new Gemini 2.5 Flash-Lite variant in preview as part of the broader Gemini 2.5 family rollout. Flash-Lite is positioned as the most cost-effective and fastest option in the Gemini 2.5 lineup, designed as a cost-effective upgrade over the p According to the article, Gemini 2.5 Flash-Lite delivers higher overall quality than 2.0 Flash-Lite in programming, mathematics, science, reasoning, and multimodal benchmarks, and is especially tuned for high-volume, latency-sensitive tasks such as translation and classification, where it achieves lower latency than 2.
This exact model name is also listed by 16 other providers.