Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
AIHubMix logo

Model details

Gemini 2.5 Flash

Gemini the listed price Flash is positioned within Google's Gemini the listed price family as the price-performance workhorse, designed for large-scale processing, low-latency interactions, and agentic workflows that benefit from built-in reasoning. Google describes it as well-rounded and particularly suited to high-volume tasks that require thinking, making it a practical choice when teams need responsive inference without stepping up to a larger Pro-tier model. Its general availability, alongside Gemini the listed price Pro, was announced as part of the broader Gemini the listed price family expansion in mid-2025.

The model accepts text, images, video, and audio as input and produces text output, supporting multimodal understanding while keeping generation simple and predictable. It enables structured outputs, function calling, code execution, search and Maps grounding, URL context, and caching, giving developers a flexible toolset for building agents and integrated pipelines. With a very large input window and a sizable output limit, Gemini the listed price Flash is well matched to long-context summarization, document and media analysis, and orchestrated agent systems where throughput and cost efficiency matter as much as raw capability.

AIHubMixgemini-2.5-flashgemini-flash

Quick Info

Powered by
Provider
AIHubMix
Model key
gemini-2.5-flash
Release date
Mar 20, 2025
Last updated
Jun 5, 2025
Knowledge cutoff
2025-01
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.30
Output token cost
$2.50

Limits

Output tokens
65,536 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare Gemini 2.5 Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Gemini 2.5 Flash

OpenRouter

Official sourceDocumentation

Google's official Gemini API documentation page for the gemini-2.5-flash model identifies it as "our best model in terms of price-performance, offering well-rounded capabilities" and recommends it for "large scale processing, low-latency, high volume tasks that require thinking, and agentic use cases." The page lists t Capabilities explicitly documented for gemini-2.5-flash include caching, code execution, file search, function calling, Grounding with Google Maps, search grounding, structured outputs, thinking, and URL context; audio generation, image generation, and the Live API are listed as not supported. Supported consumption opt

Vertex

Official sourceDocumentation

Google Cloud's official Gemini Enterprise Agent Platform documentation describes Gemini 2.5 Flash as the model's best price-to-performance option in the lineup, and notably identifies it as the first Flash model with explicit thinking capabilities, exposing the step-by-step reasoning process to developers. The page lis The same documentation page details the consumption and deployment surface available for Gemini 2.5 Flash, including Provisioned Throughput, batch inference, and pay-as-you-go options, with explicit "Not supported" flags for Gemini Live API, computer use (preview), and agentic video understanding (preview). It also poi

Videos about Gemini 2.5 Flash

More models around Gemini 2.5 Flash