Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Vercel AI Gateway logo

Model details

gpt-realtime-2

The GPT Realtime 2.x series is a speech-to-speech model line with built-in reasoning, accepting audio input and producing audio output rather than transcribing to text along the way. Azure-hosted documentation frames the family as designed for low-latency, interactive voice experiences where stronger instruction following and on-the-fly reasoning matter more than in earlier realtime models, making the architecture a fit for conversational agents, voice-driven assistants, and other scenarios that need responsive spoken exchange with the model thinking through user intent before responding.

Beyond the core audio-in, audio-out design, the 2.x preview introduces a configurable reasoning effort control that lets developers tune how much deliberation the model invests per turn, paired with response phases that separate internal commentary from the final spoken answer so downstream systems can distinguish planning from delivered output. A substantially expanded 256,000-token context window also positions the family for longer, multi-turn dialogues where earlier realtime models would have run out of room, supporting richer conversational memory and more sustained interactive sessions.

Vercel AI Gatewayopenai/gpt-realtime-2gpt

Quick Info

Powered by
Provider
Vercel AI Gateway
Model key
openai/gpt-realtime-2
Release date
May 7, 2026
Last updated
May 7, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$4.00
Output token cost
$24.00

Limits

Output tokens
0 tokens
Context window
0 tokens

Latest news about gpt-realtime-2

Vercel AI Gateway

Coverage

The Rundown's tool overview describes GPT-Realtime-2 as OpenAI's reasoning-capable speech-to-speech model for live voice agents, supporting adjustable reasoning, function calling, interruption handling, expressive delivery, and a 128K context window. The page lists inputs as audio, text, and images, outputs as audio an For developer use cases the article highlights voice-to-action agents that need reasoning and multi-step tool use, customer support calls with natural corrections and empathetic delivery, live spoken guidance that turns system context into immediate explanations, and longer voice workflows supported by the 128K context

Videos about gpt-realtime-2

More models around gpt-realtime-2