Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Kilo Gateway logo

Model details

OpenAI: GPT Audio

GPT Audio is a multimodal architecture designed to unify speech understanding and text-to-speech generation within a single call. By integrating these processes, the model serves as a primary solution for developers building short-turn voice agents that require fluid, real-time interaction. Its design intent focuses on streamlining the pipeline for audio-in and audio-out chat completions, removing the need for separate, complex systems to handle speech processing and response generation.

The model features an upgraded decoder specifically engineered to produce more natural-sounding voices while maintaining high levels of voice consistency throughout extended interactions. This advancement in the decoding process allows for more reliable performance in conversational settings. With its capacity to handle large context windows, the model is well-suited for extended dialogues, offering a robust foundation for applications that prioritize high-quality, consistent, and responsive vocal communication.

Kilo Gatewayopenai/gpt-audiogpt

Quick Info

Powered by
Provider
Kilo Gateway
Model key
openai/gpt-audio
Release date
Jan 19, 2026
Last updated
Jan 19, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$2.50
Output token cost
$10.00

Limits

Output tokens
16,384 tokens
Context window
128,000 tokens

Transparent token rates

Compare OpenAI: GPT Audio pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about OpenAI: GPT Audio

No articles yet. Fetch the latest news to show it here.

Videos about OpenAI: GPT Audio

More models around OpenAI: GPT Audio