Currently listed through these providers:
Model details
Gemini Omni Flash Preview
Gemini Omni Flash Preview is positioned by Google as a chat-oriented member of the Gemini family, designed for conversational assistants, drafting, summarization, and general question answering. Its multimodal intake lets it process text along with image and PDF attachments, while output stays in text form, making it well suited to workflows that mix document review with written responses. Within the broader Gemini lineup it is offered as a lighter, faster alternative characterized by a Flash-style profile rather than a deep reasoning variant, with typical chat usage covering routine enterprise and developer tasks that benefit from quick turnarounds and large context capacity.
In practical terms the model pairs very wide context handling with a substantial per-response output ceiling, giving teams room to ingest long reports or codebases and return detailed analyses in a single call. Access through Vercel AI Gateway is delivered via an OpenAI-compatible endpoint, so existing chat and agent frameworks that speak that protocol can route to Gemini Omni Flash Preview with minimal integration work, while keeping the underlying Google pricing tiers transparent. The combination of multimodal input, high-capacity context, and protocol-standard access makes it a pragmatic choice for production assistants, retrieval-augmented chatbots, and document-grounded Q&A where both speed and breadth of input matter.
Quick Info
Powered by- Provider
- Vercel AI Gateway
- Model key
- google/gemini-omni-flash-preview
- Release date
- Jun 30, 2026
- Last updated
- Jun 30, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $1.50
- Output token cost
- $9.00
Limits
- Output tokens
- 57,920 tokens
- Context window
- 1,000,000 tokens