Currently listed through these providers:
Model details
Gemini 3.5 Flash
Gemini 3.5 Flash builds on the Gemini 3 Flash reasoning foundation by adding explicit thinking levels that let developers tune the balance between quality, cost, and response speed. Google positioned this model as delivering Pro-level reasoning at Flash-class latency, with the aim of making agentic and coding workloads once reserved for higher-tier models accessible at lower operational costs. The architecture supports function calling, structured output, code execution, and search-as-tool directly from first-party tooling, making it practical for developers working across diverse service categories.
Introduced at Google I/O in May 2026 as part of the broader Gemini 3.5 family, the model went through benchmark evaluation via Google's published model card and independent testing platforms. Independent assessments across 191 questions spanning nine categories found it capable enough to handle complex agentic tasks. The model carries a January 2025 knowledge cutoff, keeping its training data current enough for most production use cases while still demonstrating strong performance on practical coding and reasoning challenges.
Quick Info
Powered by- Provider
- Venice AI
- Model key
- gemini-3-5-flash
- Release date
- May 22, 2026
- Last updated
- Jun 11, 2026
- Knowledge cutoff
- 2025-01
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $1.55
- Output token cost
- $9.45
Limits
- Output tokens
- 65,536 tokens
- Context window
- 1,000,000 tokens