Poe
GPT-5.3-Codex moved to No. 1 in Quality on the Microsoft Foundry AI Model Leaderboard soon after release, while a cross-metric "podium" scoring method put GPT-5-Nano on top overall for efficiency.
Model details
GPT-5.3-Codex is designed as a highly capable agentic model that expands the scope of automated software development beyond simple code generation and review. By integrating advanced reasoning with professional knowledge, the model is built to function as a collaborative colleague capable of executing long-running tasks that involve research, tool usage, and complex system interactions. It is engineered to maintain context during these extended workflows, allowing users to provide ongoing steering and supervision. This design intent positions the model to handle nearly any task a professional developer might perform on a computer, from managing codebase coherence to executing deep, multi-step engineering projects.
The development of this model represents a unique milestone in AI lineage, as it is the first system to be instrumental in its own creation. Early versions were utilized by the development team to debug training processes, manage deployment, and diagnose evaluation results, effectively accelerating the model's own evolution. This iterative approach has resulted in a system that delivers significant performance gains on industry benchmarks like SWE-Bench Pro, Terminal-Bench 2.0, and OSWorld-Verified. Beyond its core coding strengths, the model is optimized for practical, real-world utility, offering improved transparency through deep diffs and better handling of complex logic, making it a robust tool for teams engaged in automated pull requests and issue-to-patch workflows.
Poe
GPT-5.3-Codex moved to No. 1 in Quality on the Microsoft Foundry AI Model Leaderboard soon after release, while a cross-metric "podium" scoring method put GPT-5-Nano on top overall for efficiency.
Poe
GPT-5.3-Codex, OpenAI's latest agentic coding model, is now rolling out in GitHub Copilot. In early testing, GPT-5.3-Codex reaches new high scores on...
Poe
BenchLM's tracker entry for GPT-5.3 Codex, data as of September 3, 2026, lists the model as released February 5, 2026, with a 400K token context window, an aggregate capability score of 65.5 out of 100, and a rank of 42 of 231 tracked models. It reports API pricing of $1.75 per million input tokens and $14 per million The same BenchLM page publishes category-level percentiles for GPT-5.3 Codex against its eligible cohort: 86th percentile in Agentic (rank 21 of 144) and 83rd percentile in Coding (rank 26 of 149), while Reasoning, Knowledge, Math, Multilingual, Multimodal, and Instruction Following are not ranked. The tracker notes th