Currently listed through these providers:
Model details
Gemini 3.5 Flash
Gemini 3.5 Flash arrives as a mid-tier reasoning model that Google pitches as delivering "Pro-level reasoning at Flash-class latency," aiming to absorb agentic and coding workloads that previously lived on the Pro tier. Internally, it is built on the Gemini 3 Flash reasoning foundation and exposes explicit thinking-level controls that let developers trade quality against latency and cost on a per-request basis. Third-party reviewers have taken that positioning seriously and put it through three reference lenses: Google's own model card, the Artificial Analysis leaderboard entry, and Appwrite's open-source Arena benchmark of 191 questions across nine Appwrite service categories.
Because the catalog listing routes this model through AIHubMix, it fits naturally into teams already standardising on that gateway's unified, OpenAI-compatible Chat Completions endpoint, where it sits alongside other Gemini-family entries. Practical fit is best in high-volume, latency-sensitive pipelines where the model can carry routine reasoning, tool orchestration, and code-assistance duties without needing a heavier Pro-tier deployment. Teams picking between tiers can use the Appwrite and Artificial Analysis comparisons as a grounded starting point before committing production traffic.
Quick Info
Powered by- Provider
- AIHubMix
- Model key
- gemini-3.5-flash
- Release date
- May 19, 2026
- Last updated
- May 19, 2026
- Knowledge cutoff
- 2025-01
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $1.50
- Output token cost
- $9.00
Limits
- Output tokens
- 64,000 tokens
- Context window
- 1,000,000 tokens
Latest news about Gemini 3.5 Flash
Videos about Gemini 3.5 Flash
More models around Gemini 3.5 Flash
This exact model name is also listed by 21 other providers.