Currently listed through these providers:
Model details
GPT Audio Mini
GPT Audio Mini sits inside the broader GPT Audio family as a lighter, more affordable option aimed at developers who want speech capabilities without stepping up to the flagship tier. Its identity as a value-oriented variant is clear from how OpenRouter introduces it, framing it as a cost-efficient version of GPT Audio that still delivers conversational voice output. The model accepts both text and audio on the input side and can respond in either text or audio, making it suitable for mixed-modality assistants where users may type, speak, or listen within the same session. The latest snapshot focuses on improving the listening experience rather than chasing raw capability ceilings. OpenRouter notes that the release brings an upgraded decoder for more natural-sounding voices and improved voice consistency across longer exchanges, which matters for applications like voice agents, audiobooks, and interactive tutoring where prosody and speaker stability shape perceived quality. Because it carries the same modalities as the broader GPT Audio line at a lower price point, GPT Audio Mini is a natural fit for production voice features where cost per interaction, predictable latency, and consistent vocal character matter more than maximum reasoning depth.
From a deployment standpoint, the model is positioned for straightforward integration through a single upstream provider, with OpenRouter routing requests directly to OpenAI rather than balancing across multiple hosts. That single-provider routing keeps latency predictable and simplifies observability, since there is no abstraction layer choosing among backends for each call. Developers pairing it with prompt caching or repeated system prompts can see effective per-token costs well below the listed rates, which broadens the range of interactive voice use cases where the economics work out. For teams deciding between this model and the standard GPT Audio option, the trade-off is essentially quality versus scale: GPT Audio Mini trades a degree of top-end capability for meaningfully cheaper inference and the same multimodal surface area. That makes it well suited to high-volume customer-facing features such as multilingual support lines, in-app voice companions, and content narration pipelines where natural-sounding speech and reliable voice identity are the primary requirements. Teams that need deeper reasoning or longer-form generation can still reach for the larger sibling, while everyday voice workloads can lean on GPT Audio Mini as a dependable, budget-friendly default.
Quick Info
Powered by- Provider
- OpenRouter
- Model key
- openai/gpt-audio-mini
- Release date
- Jan 19, 2026
- Last updated
- Jan 19, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.60
- Output token cost
- $2.40
Limits
- Output tokens
- 16,384 tokens
- Context window
- 128,000 tokens