Vertex
Google has added agentic vision to Gemini 3 Flash, combining visual reasoning with code execution to "ground answers in visual evidence". According to Google, this not only improves accuracy, but more
Model details
Gemini the listed price Flash Preview sits in the Gemini Flash family as a speed- and value-oriented thinking variant intended for agentic workflows, multi-turn chat, and coding assistance. It is positioned as delivering reasoning and tool-use quality close to the larger Gemini Pro models while keeping latency substantially lower, making it well suited to interactive development, long-running agent loops, and collaborative coding sessions where responsiveness matters as much as raw capability. Compared with its predecessor Gemini 2.5 Flash, it is described as offering broad quality improvements across reasoning, multimodal understanding, and reliability, reflecting a generational step rather than a narrow tuning change.
In practical terms, Gemini the listed price Flash Preview is built for developers who want strong reasoning and agentic behavior without paying for frontier-scale latency. It supports a one-million-token context window and accepts text, image, audio, video, and PDF inputs while producing text output, which lets a single model handle mixed-media grounding, long document reasoning, and tool-augmented chat in one pass. Developers can tune the depth of inference through configurable thinking levels (minimal, low, medium, high), and the model exposes structured output, tool use, and automatic context caching so that agent pipelines can keep working state, call external functions, and reuse prior context efficiently across multi-step tasks.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Vertex
Google has added agentic vision to Gemini 3 Flash, combining visual reasoning with code execution to "ground answers in visual evidence". According to Google, this not only improves accuracy, but more
Vertex
Gemini 3 Flash Preview API pricing: $0.5/1M input, $3/1M output. 1M context. Supports function calling, web search, reasoning. Access via Inworld Router with OpenAI SDK compatibility and failover.