Currently listed through these providers:
Model details
OpenAI: GPT Audio
GPT Audio is a multimodal architecture designed to unify speech understanding and text-to-speech generation within a single call. By integrating these processes, the model serves as a primary solution for developers building short-turn voice agents that require fluid, real-time interaction. Its design intent focuses on streamlining the pipeline for audio-in and audio-out chat completions, removing the need for separate, complex systems to handle speech processing and response generation.
The model features an upgraded decoder specifically engineered to produce more natural-sounding voices while maintaining high levels of voice consistency throughout extended interactions. This advancement in the decoding process allows for more reliable performance in conversational settings. With its capacity to handle large context windows, the model is well-suited for extended dialogues, offering a robust foundation for applications that prioritize high-quality, consistent, and responsive vocal communication.
Quick Info
Powered by- Provider
- Kilo Gateway
- Model key
- openai/gpt-audio
- Release date
- Jan 19, 2026
- Last updated
- Jan 19, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $2.50
- Output token cost
- $10.00
Limits
- Output tokens
- 16,384 tokens
- Context window
- 128,000 tokens
Transparent token rates
Compare OpenAI: GPT Audio pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about OpenAI: GPT Audio
No articles yet. Fetch the latest news to show it here.