Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
OpenAI logo

Model details

GPT-4o mini

GPT-4o mini was introduced by OpenAI as a compact derivative of the larger GPT-4o system, shaped through a distillation-style process that aims to preserve the parent model's capabilities while sharply reducing cost and latency. In its launch announcement, OpenAI positioned the model as its most cost-efficient small option, designed to make high-quality AI broadly accessible for developers and businesses that need to embed intelligence into everyday products rather than run expensive frontier-scale inference. The design philosophy centers on small-model economics combined with the upgraded tokenizer inherited from GPT-4o, which improves handling of non-English text and keeps per-token costs low across languages.

In practical terms, GPT-4o mini was built for high-throughput, long-context workflows such as chaining many model calls in parallel, feeding in full code bases or extended conversation histories, and serving fast real-time text responses like customer support chatbots. OpenAI reported that the model achieved an 82% score on MMLU at launch and, on its announcement day, outperformed GPT-4 on chat-preference rankings in the LMSYS chatbot arena, signaling competitive reasoning quality for its size class. This combination of strong benchmark performance, a wide context span, and pricing well below earlier OpenAI offerings makes the model a practical fit for cost-sensitive production deployments where responsiveness and scale matter more than maximum capability.

OpenAIgpt-4o-minigpt-mini

Quick Info

Powered by
Provider
OpenAI
Model key
gpt-4o-mini
Release date
Jul 18, 2024
Last updated
Jul 18, 2024
Knowledge cutoff
2023-09
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$0.60

Limits

Output tokens
16,384 tokens
Context window
128,000 tokens

Transparent token rates

Compare GPT-4o mini pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GPT-4o mini

OpenAI

Official sourceRelease Notes

The OpenAI API changelog from August 2026 announces the deprecation of gpt-4o-mini-transcribe alongside whisper-1, gpt-4o-transcribe, and gpt-4o-transcribe-diarize, with a scheduled shutdown date of February 26, 2027. OpenAI directs affected developers to migrate to gpt-live-transcribe or gpt-transcribe and links to it An earlier August 2026 changelog entry also introduced mutual TLS (mTLS) and X.509 workload identity federation as generally available for the OpenAI API, allowing certificate and identity-provider configuration directly in the Platform console under organizational roles and permissions. A separate August 21, 2026 upda

OpenAI

CoverageBenchmark

GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. $0.15 per million input tokens, $0.60 per million output tokens. 128,000 token context window, maximum output of 16,384 tokens. Higher uptime with 2 providers. Includes independent be

Videos about GPT-4o mini

More models around GPT-4o mini