Currently listed through these providers:
Model details
GPT-Realtime-2.1
GPT-Realtime-2.1 is positioned as a next-generation speech-to-speech reasoning model purpose-built for live conversational interactions. It processes streaming audio input and produces audio responses while carrying out built-in reasoning as part of the response pipeline, rather than treating reasoning as a separate stage. The design emphasis on continuous, low-latency streaming makes it well suited for voice assistants, interactive agents, and other scenarios where responsiveness during an ongoing dialogue matters more than batch processing.
On Microsoft Foundry the model is listed as a Direct from Azure offering under the Azure OpenAI service, sitting within a curated portfolio that Microsoft frames as secure and managed, with unified billing, governance, and portable reserved capacity across hosted models. This commercial positioning points to enterprise-oriented deployment rather than self-hosted open weights, and the Direct from Azure framing suggests the model is meant to be adopted alongside other Azure-hosted voice and reasoning models for consistent operations. For practitioners, it offers a path to streaming voice experiences with embedded reasoning on managed infrastructure, with the practical fit being real-time customer-facing voice products that benefit from a single vendor-managed stack.
Quick Info
Powered by- Provider
- OpenAI
- Model key
- gpt-realtime-2.1
- Release date
- Jul 6, 2026
- Last updated
- Jul 6, 2026
- Knowledge cutoff
- 2024-09-30
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $4.00
- Output token cost
- $24.00
Limits
- Input tokens
- 96,000 tokens
- Output tokens
- 32,000 tokens
- Context window
- 128,000 tokens