Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
FastRouter logo

Model details

GPT Realtime 1.5

GPT Realtime 1.5 is an OpenAI voice and audio model documented in the company's API model catalog under the Audio and voice section, intended for building real-time conversational agents that can listen, understand, and respond in natural spoken language. Third-party coverage describes it as a targeted voice AI upgrade, emphasizing its ability to maintain flowing dialogue while invoking external tools such as calendars, databases, or booking APIs without breaking the conversational rhythm. The same reporting highlights multilingual handling, including switching languages mid-sentence, positioning the model for global customer service, accessibility, and live-assistant use cases where fluid spoken interaction matters more than long-form reasoning.

The reported advances in GPT Realtime 1.5 center on audio-specific reasoning and transcription quality rather than general-purpose text performance. Community coverage attributes measurable gains to the release, including a roughly five percent improvement on audio reasoning with Big Bench Audio reaching 82.8 percent, a ten-point gain in accuracy when transcribing numbers, codes, and alphanumeric strings, and a seven percent lift in instruction following during live conversations. These gains, paired with mid-conversation tool calling and multilingual code-switching, make the model a practical fit for developers building voice-driven assistants that must reliably parse precise identifiers and act on user intent in real time.

FastRouteropenai/gpt-realtime-1.5gpt

Quick Info

Powered by
Provider
FastRouter
Model key
openai/gpt-realtime-1.5
Release date
Jun 1, 2025
Last updated
Jun 1, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$4.00
Output token cost
$16.00

Limits

Output tokens
4,096 tokens
Context window
32,000 tokens

Transparent token rates

Compare GPT Realtime 1.5 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GPT Realtime 1.5

FastRouter

CoverageBenchmark

Synthszr's product-ranking page explicitly names "GPT-Realtime-1.5" and identifies it as "OpenAI's flagship real-time audio model for voice agents and customer support," describing it as improving on its predecessor with "better instruction following, multilingual handling, and tool calling while preserving low-latency The third-party page adds a latency data point: median turn latency of 2.24 seconds on medium-length calls (30 turns), rising to 3.4 seconds median at 60 turns, which is a useful operational benchmark for sizing realtime voice-agent sessions. It also notes "seamless mid-sentence language switching" as an explicitly imp

FastRouter

CoverageRelease Notes

Microsoft's "What's new in Azure OpenAI in Microsoft Foundry Models" page documents, under a February 2026 release section, that "the gpt-realtime-1.5 and gpt-audio-1.5 models are now available" for Azure customers. The release note states these models "build on last year's GPT-Realtime and GPT-Audio with improvements For developers building voice-first applications on Azure, the February 2026 release note is significant because gpt-realtime-1.5 is exposed via the existing Chat Completions API rather than requiring a separate realtime client, lowering the integration cost for teams already on Azure OpenAI. The explicit mention of im

Videos about GPT Realtime 1.5

More models around GPT Realtime 1.5