Currently listed through these providers:
Model details
Gemini 3.5 Flash Lite (Google Vertex AI)
Gemini 3.5 Flash-Lite is positioned for fast, economical execution where many requests must be handled quickly. Its intended workloads include document parsing, straightforward data extraction, and subagent or other high-volume agentic tasks, with multimodal inputs available when an application needs to work across text, images, video, audio, or PDF files. The source describes it as proprietary, but does not provide a parameter count, architecture details, training scale, or benchmark results.
In practice, the model suits pipelines in which response speed and operating cost matter more than a claim of frontier leadership. Its listing supports function calling, structured output, reasoning, JSON mode, streaming, fine-tuning, and batch processing, making it relevant for extraction services, document workflows, and coordinated agents. Those capabilities and the large context should be treated as catalog metadata from a third-party hub rather than a substitute for Google-official technical documentation.
Quick Info
Powered by- Provider
- LLM Gateway
- Model key
- google-vertex/gemini-3.5-flash-lite
- Release date
- Jul 21, 2026
- Last updated
- Jul 21, 2026
- Knowledge cutoff
- 2026-03
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.30
- Output token cost
- $2.50
Limits
- Output tokens
- 65,536 tokens
- Context window
- 1,048,576 tokens
Transparent token rates
Compare Gemini 3.5 Flash Lite (Google Vertex AI) pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Gemini 3.5 Flash Lite (Google Vertex AI)
No articles yet. Fetch the latest news to show it here.
Videos about Gemini 3.5 Flash Lite (Google Vertex AI)
More models around Gemini 3.5 Flash Lite (Google Vertex AI)
This exact model name is also listed by 24 other providers.
