Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
ZenMux logo

Model details

GPT-5.1-Codex

GPT-5.1-Codex is engineered as an agentic coding companion built for long-running, multi-step development workflows rather than simple code completion. It introduces a context compaction technique that lets it work coherently across multiple context windows, effectively handling million-token codebases in a single task. The model accepts both text and image inputs, enabling it to reason across code, UI states, architecture diagrams, and design comps within the same workflow. Its repo-aware intelligence understands full repositories, supporting cohesive refactors and test automation, while model-guided loops retain state and context across extended interactions for truly asynchronous execution of long-running coding tasks.

Post-training incorporates reinforcement learning and supervised fine-tuning to cultivate agentic behavior, with reasoning effort levels that let developers trade latency for code quality depending on task complexity. The xhigh reasoning level achieved 77.9% on SWE-bench Verified while using 30% fewer thinking tokens, and Terminal Bench 2.0 scores of 58.1% outpaced competing models from Gemini and Anthropic. An independent METR evaluation found the model posed low risk for AI R&D automation and rogue replication threats, with OpenAI observing the system operate autonomously for over 24 hours continuously—iterating through code and fixing failures without human intervention. These capabilities make it well-suited for teams deploying autonomous development agents at scale.

ZenMuxopenai/gpt-5.1-codex

Quick Info

Powered by
Provider
ZenMux
Model key
openai/gpt-5.1-codex
Release date
Nov 13, 2025
Last updated
Nov 13, 2025
Knowledge cutoff
2025-01-01
AI SDK package
@ai-sdk/openai
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$1.25
Output token cost
$10.00

Limits

Output tokens
64,000 tokens
Context window
400,000 tokens

Latest news about GPT-5.1-Codex

Vercel AI Gateway

CoverageBenchmark

BenchLM's GPT-5.1-Codex page, published roughly five days before the current date, summarizes third-party benchmark evidence for the model. It assigns an overall capability score of 48.7/100 based on a field median of 56.4 across 127 of 231 ranked models, alongside pricing of $1.25 input and $10 output per 1M tokens an The aggregator reports verified benchmark rows only in the Agentic and Coding categories — Agentic rank 65 of 154 (58th percentile, 2 benchmarks) and Coding rank 88 of 156 (44th percentile, 2 benchmarks). Reasoning, Multimodal, Knowledge, Multilingual, Instruction-Following, and Math categories are all marked as Not me

Videos about GPT-5.1-Codex