Anthropic introduced Claude Opus 4.6 as an upgrade to its flagship Opus line, sharpening the model for software engineering and long-running agentic work. The company describes more careful planning, sustained execution on multi-step tasks, more reliable behavior inside large codebases, and improved code review and debugging that helps the model catch its own mistakes. Anthropic also highlights the model's ability to handle everyday professional workflows such as financial analyses, research, and working with documents, spreadsheets, and presentations, with autonomous multitasking available through its Cowork research preview.
On evaluations, Anthropic reports that Claude Opus 4.6 sets a new high on the Terminal-Bench 2.0 agentic coding benchmark and leads other frontier models on Humanity's Last Exam, a multidisciplinary reasoning test. On GDPval-AA, which measures economically valuable knowledge work, Anthropic claims Opus 4.6 outperforms the next-best competitor by roughly 144 Elo points and its own predecessor by 190 points, and it also tops BrowseComp for locating hard-to-find information. A standout architectural change is the introduction of a one-the cataloged API limit for an Opus-class model, offered in beta, giving the system room to reason over very large inputs such as full codebases or lengthy document collections.