Currently listed through these providers:
Model details
GPT-4o
GPT-4o, where the "o" stands for omni, was designed as a step toward much more natural human-computer interaction by accepting any combination of text, audio, image, and video as input and generating outputs across text, audio, and image formats. The model can respond to audio inputs in as little as 232 milliseconds, with an average of 320 milliseconds—response times similar to human conversation pace. It was built to especially excel at vision and audio understanding compared to existing models, while also delivering significant improvements in processing non-English languages.
The model matches GPT-4 Turbo performance on English text and code tasks while being twice as fast and 50% more cost-effective in API usage. While current API access supports text and image inputs with text outputs (the same modalities as GPT-4 Turbo), audio and additional output modalities are planned for future introduction. This combination of speed, cost efficiency, and multimodal capabilities positions GPT-4o as a practical upgrade for developers building applications that demand fast, responsive interactions across different content types.
Quick Info
Powered by- Provider
- OpenRouter
- Model key
- openai/gpt-4o
- Release date
- May 13, 2024
- Last updated
- Aug 6, 2024
- Knowledge cutoff
- 2023-09
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $2.50
- Output token cost
- $10.00
Limits
- Output tokens
- 16,384 tokens
- Context window
- 128,000 tokens