OpenRouter
Compare GPT-5.1-Codex-Max from OpenAI to other AI models on key metrics including benchmarks, price, context length, and other model features.
Model details
GPT-5.1-Codex-Max is positioned as a frontier agentic coding model built upon a revamped reasoning architecture, specifically engineered to handle long-running and complex development workflows. Its defining innovation is a technique called compaction, which enables the model to operate coherently across multiple context windows and work with millions of tokens within a single session. This design choice directly addresses the limitations of traditional models in sustained development tasks, unlocking capabilities such as project-scale refactoring, deep debugging sessions spanning hours, and multi-hour agent loops that can run continuously for over a day. The architecture was trained on agentic tasks spanning software engineering, mathematics, and research, with particular attention to real-world workflows like pull request creation, code review, and frontend development. Notably, this model marks the first time a model was natively trained to operate within Windows environments, reflecting a deliberate push toward broader development environment compatibility.
The training pipeline leveraged realistic software engineering tasks designed to make the model a more effective collaborator in the Codex CLI environment. Improved token efficiency means that even at medium reasoning effort, the model achieves better results than previous generations using fewer tokens, with an option to select extra high reasoning effort when quality matters more than latency. Evaluation benchmarks like Terminal-Bench2.0 were run with compaction enabled at extra high reasoning effort, and the model outperformed prior iterations across many frontier coding evaluations. Independent assessment by METR found the model represents a low-risk incremental improvement over earlier thinking models, with no risk-critical changes to architecture or training incentives. The practical strengths of this model make it well-suited for development teams seeking an AI partner capable of handling extended, multi-step engineering tasks without losing context or requiring constant intervention.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
OpenRouter
Compare GPT-5.1-Codex-Max from OpenAI to other AI models on key metrics including benchmarks, price, context length, and other model features.
OpenRouter
Compare GPT-5.1-Codex-Max from OpenAI and Qwen3.5 Plus 2026-02-15 from Qwen on key metrics including benchmarks, price, context length, and other model features.
This exact model name is also listed by 10 other providers.