Gemini 3.5 Flash-Lite is positioned within Google's 3.x Flash family as the lightweight counterpart that pairs low latency with cost-effective throughput for production workloads. According to the official Gemini API changelog, it reached general availability on July 21, 2026 alongside Gemini 3.6 Flash, as part of a stable, production-ready release of the latest Flash-tier lineup. The model is explicitly framed as "a low-latency, highly cost-effective subagent option designed for high-volume automation," which signals that Google optimized the Flash-Lite branch for scaling many parallel, focused tasks rather than for single-shot deep reasoning.
In practice, the variant is described as a high-efficiency model with upgraded agentic capabilities, well suited to subagents that handle discrete responsibilities inside larger multi-agent systems. Vertex AI listings confirm a one-the cataloged API limit and a roughly 66,000-token maximum output, giving subagents enough room to consume long tool traces, retrieval snippets, or code context while still returning substantial structured responses. The capability surface—spanning vision, reasoning, tool calling, caching, web search, and JSON-schema structured output—makes it flexible enough to plug into orchestrator-driven pipelines where one Flash-Lite instance might classify, route, or summarize before handing richer generation off to a larger model in the same agent graph.