Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Requesty logo

Model details

Gemini 2.5 Flash-Lite (EU)

Gemini 2.5 Flash-Lite (EU) sits at the lightweight end of Google's 2.5 generation, designed as a low-cost lane for high-throughput multimodal traffic and quick agent loops. It carries the same Gemini 2.5 lineage as the larger Flash and Pro variants but is positioned for lean workloads where input cost and responsiveness matter more than deep reasoning. The EU routing through Requesty gives European teams a familiar procurement path while keeping the broader Gemini multimodal surface area intact.

In practical terms, the model suits production use cases that demand cheap, broad-coverage input handling rather than frontier reasoning: short-form chat, routing agents, summarization, classification, extraction, and lightweight code or translation passes. Its 1M-token context window is unusually large for a Flash-Lite tier, making it useful for long-document digestion and multi-turn tool pipelines where the Flash or Pro tiers would be overkill. Teams building high-volume assistants or background automation get a straightforward upgrade path within the Gemini family without paying for Pro-class depth.

Requestygemini-2.5-flash-lite@eugemini-flash-lite

Quick Info

Powered by
Provider
Requesty
Model key
gemini-2.5-flash-lite@eu
Release date
Jun 17, 2025
Last updated
Jun 17, 2025
Knowledge cutoff
2025-01
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.10
Output token cost
$0.40

Limits

Output tokens
65,535 tokens
Context window
1,048,576 tokens

Latest news about Gemini 2.5 Flash-Lite (EU)

Requesty

Official sourceBenchmark

Requesty documents a Vertex AI-hosted deployment of Google's Gemini 2.5 Flash-Lite pinned to europe-west1, exposing it through the OpenAI-compatible endpoint at https://router.requesty.ai/v1 with model id vertex/gemini-2.5-flash-lite@europe-west1. The page describes this as an EU-only, no-routing/no-failover deployment The model offers a 1.0M token context window with 66K output and capability flags for Vision, Reasoning, Tool calling, Caching, Web search, JSON schema, Computer use, and Image generation. Live production telemetry from Requesty traffic reports a 398ms median time-to-first-token, 233 tokens-per-second median output spe

Videos about Gemini 2.5 Flash-Lite (EU)

More models around Gemini 2.5 Flash-Lite (EU)