Sulat.com
AI models
Azure logo

Model details

GPT-4.1 nano

GPT-4.1 nano arrived as part of the GPT-4.1 family launch and was described as OpenAI's first nano-tier offering, positioned as a lightweight sibling to the standard and mini variants. It inherits the family-level design priorities of improved instruction following, stronger coding ability, and much larger context handling, with the family supporting up to one million tokens of context overall. The model carries a refreshed knowledge cutoff of June 2024 and accepts both text and image inputs while producing text-only output, making it suitable for multimodal pipelines where cost and speed matter more than frontier reasoning quality.

From a practical standpoint, GPT-4.1 nano is aimed at high-volume, latency-sensitive workloads such as classification, extraction, routing, and short-form generation that benefit from a million-token context without paying full-size model prices. Independent benchmarking on Azure-hosting reports solid quality for its tier, with GPQA Diamond around 49 percent and Tau-Bench near 14.7 percent, while output throughput reaches roughly 241 tokens per second on Azure, making it one of the faster small models available. Because the model is now deprecated with a recommendation to migrate to newer nano-tier successors, it is best treated as an interim option for existing applications rather than a starting point for new builds.

Azuregpt-4.1-nanogpt-nanodeprecated

Quick Info

Powered by
Provider
Azure
Model key
gpt-4.1-nano
Release date
Apr 14, 2025
Last updated
Apr 14, 2025
Knowledge cutoff
2024-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.10
Output token cost
$0.40

Limits

Output tokens
32,768 tokens
Context window
1,047,576 tokens

Transparent token rates

Compare gpt-nano pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GPT-4.1 nano

Azure

CoverageBenchmark

Artificial Analysis notes that GPT-4.1 nano is now deprecated, with OpenAI recommending migration to GPT-5 nano (medium), and the page only continues benchmarking the default 10k input token workload. Despite deprecation, it reports Azure as the fastest provider for output speed at 241.4 tokens per second, compared to The comparison covers cache-hit, input, and output pricing for both Azure and OpenAI at the 10,000 input token workload, with separate notes on cache write and storage billing that vary by provider. The analysis is intended to help developers choose between the two GPT-4.1 nano API providers based on throughput, latenc

Azure

CoverageBenchmark

OpenRouter's listing for GPT-4.1 Nano confirms Azure is one of two hosted providers alongside OpenAI, with Azure listed at $0.10 input / $0.40 output per 1M tokens, a cache-read price of $0.03, a measured P50 latency of 0.73s, throughput of 15 tps, and 99.98% uptime. The model supports a 1M token context window, scores The page also reports per-provider benchmark performance: Azure scores 49.0% on GPQA Diamond and 14.7% on Tau-Bench, compared to OpenAI's 47.1% GPQA and 14.0% Tau-Bench scores, while auto-routing across providers yields 51.9% GPQA and 10.7% Tau-Bench. OpenRouter's three-day availability for GPT-4.1 Nano reaches 100% up

Videos about GPT-4.1 nano

More models around GPT-4.1 nano