Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
302.AI logo

Model details

gpt-4.1-nano

GPT-4.1-nano belongs to the same family line as GPT-4.1 and GPT-4.1-mini, introduced as the smallest and most economical member of that series. It inherits the GPT-4.1 family's core strengths, including multimodal input handling for both text and images, long-context understanding, improved instruction following, and enhanced coding capabilities. Its design goal is to deliver that capability set at substantially reduced latency and cost, making it well suited for real-time applications and large-scale deployment of parallel or chained model calls, such as powering chatbots, coding copilots, and AI agents that need to invoke multiple lightweight models quickly.

Benchmark evidence positions GPT-4.1-nano as surprisingly capable for its size, scoring 80.1% on MMLU, 50.3% on GPQA, and 9.8% on Aider polyglot coding, which exceeds the performance of GPT-4o mini. The model operates within a one-million-token context window and supports structured output generation alongside parallel function calling for tool-based workflows. These characteristics make it a strong practical fit for lightweight, high-frequency workloads like classification, autocompletion, and other behind-the-scenes tasks where speed and cost matter more than top-tier reasoning depth.

302.AIgpt-4.1-nanogpt-nano

Quick Info

Powered by
Provider
302.AI
Model key
gpt-4.1-nano
Release date
Apr 14, 2025
Last updated
Apr 14, 2025
Knowledge cutoff
2024-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.10
Output token cost
$0.40

Limits

Output tokens
32,768 tokens
Context window
1,047,576 tokens

Transparent token rates

Compare gpt-4.1-nano pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about gpt-4.1-nano

302.AI

CoverageBenchmark

The OpenRouter aggregator listing for OpenAI's GPT-4.1-nano reports a 1 million token context window, a knowledge cutoff of June 2024, and an April 14, 2025 release date, with list pricing of $0.10 per 1M input tokens and $0.40 per 1M output tokens. Benchmark figures quoted on the page include 80.1% on MMLU, 50.3% on G Rolling 7- to 30-day performance data on the same page indicates p50 latency around 0.49 seconds and throughput up to 85 tokens per second for the best-performing host, with weighted-average effective prices of roughly $0.080 per 1M input and $0.398 per 1M output after prompt caching. Tool-call error rates are reported

Videos about gpt-4.1-nano

More models around gpt-4.1-nano