Azure
Tech News News: OpenAI has launched two new AI models — GPT-5.4 mini and GPT-5.4 nano — aimed at delivering faster performance and lower costs for high-volume workloa.
Model details
GPT-5.4 Nano sits at the small end of the GPT-5.4 family, positioned as OpenAI's lightest variant for ultra-low-latency and cost-sensitive workloads. Introduced alongside GPT-5.4 mini in March 2026 as faster, efficient siblings of the main GPT-5.4 line, it is designed to inherit many of that family's strengths while trimming the compute and price bill for high-throughput traffic. The variant is aimed at routing layers, classification, extraction, and other lightweight agent or automation tasks where responsiveness matters more than top-tier reasoning depth. It also accepts image input alongside text, broadening its use beyond pure text pipelines while still producing only text output.
For teams building production systems, GPT-5.4 Nano's practical appeal is the combination of a 400,000-token context window with pricing tuned for volume, making it suitable for backend scoring, triage, and pre-processing stages that sit in front of larger models. The model is available through Azure, where it has rolled out for Global Standard deployments, and news coverage indicates it was part of a broader expansion aimed at replacing older low-latency options reaching retirement. Its blend of multimodal input, structured tool use, and aggressive cost-efficiency makes it a natural fit for orchestration layers that need to call an inexpensive reasoning model thousands of times per minute without sacrificing the broader capabilities of the GPT-5.4 family.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Azure
Tech News News: OpenAI has launched two new AI models — GPT-5.4 mini and GPT-5.4 nano — aimed at delivering faster performance and lower costs for high-volume workloa.
Azure
On March 17, 2026, OpenAI announced the release of ' GPT-5.4 mini ' and ' GPT-5.4 nano ,' lightweight versions of GPT-5.4, which debuted in March 2026. GPT-5.4 mini/nano are designed to be fast and efficient models that can handle high processing loads while inheriting many of the strengths of GPT-5.4. OpenAI introduce
Azure
OpenRouter's GPT-5.4 Nano listing confirms Azure as a hosted provider for the model, alongside OpenAI direct, with a list price of $0.20 per 1M input tokens and $1.25 per 1M output tokens on a 400K context window. The page reports a March 17, 2026 release date and an August 2025 knowledge cutoff, and notes that GPT-5.4 Provider-level telemetry on the page shows Azure delivering GPT-5.4 Nano at roughly 1.25s P50 latency and 35 t/s throughput with 99.99% uptime, versus 0.72s and 69 t/s on OpenAI direct. Effective pricing after prompt caching averages about $0.089 per 1M input tokens across both providers, with a 61–62% cache hit rate,
This exact model name is also listed by 29 other providers.