Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
LLM Gateway logo

Model details

Gemini 3.5 Flash Lite (Google Vertex AI)

Gemini 3.5 Flash-Lite is positioned for fast, economical execution where many requests must be handled quickly. Its intended workloads include document parsing, straightforward data extraction, and subagent or other high-volume agentic tasks, with multimodal inputs available when an application needs to work across text, images, video, audio, or PDF files. The source describes it as proprietary, but does not provide a parameter count, architecture details, training scale, or benchmark results.

In practice, the model suits pipelines in which response speed and operating cost matter more than a claim of frontier leadership. Its listing supports function calling, structured output, reasoning, JSON mode, streaming, fine-tuning, and batch processing, making it relevant for extraction services, document workflows, and coordinated agents. Those capabilities and the large context should be treated as catalog metadata from a third-party hub rather than a substitute for Google-official technical documentation.

LLM Gatewaygoogle-vertex/gemini-3.5-flash-litegemini-flash-lite

Quick Info

Powered by
Provider
LLM Gateway
Model key
google-vertex/gemini-3.5-flash-lite
Release date
Jul 21, 2026
Last updated
Jul 21, 2026
Knowledge cutoff
2026-03
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.30
Output token cost
$2.50

Limits

Output tokens
65,536 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare Gemini 3.5 Flash Lite (Google Vertex AI) pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Gemini 3.5 Flash Lite (Google Vertex AI)

No articles yet. Fetch the latest news to show it here.

Videos about Gemini 3.5 Flash Lite (Google Vertex AI)

More models around Gemini 3.5 Flash Lite (Google Vertex AI)