Requesty
Requesty documents a Vertex AI-hosted deployment of Google's Gemini 2.5 Flash-Lite pinned to europe-west1, exposing it through the OpenAI-compatible endpoint at https://router.requesty.ai/v1 with model id vertex/gemini-2.5-flash-lite@europe-west1. The page describes this as an EU-only, no-routing/no-failover deployment The model offers a 1.0M token context window with 66K output and capability flags for Vision, Reasoning, Tool calling, Caching, Web search, JSON schema, Computer use, and Image generation. Live production telemetry from Requesty traffic reports a 398ms median time-to-first-token, 233 tokens-per-second median output spe