Azure Cognitive Services
Delegate complex tasks end-to-end and execute reliably with Claude Opus 4.6, now available in Microsoft Foundry.
Model details
Anthropic introduced Claude Opus 4.6 as an upgrade to its flagship Opus line, sharpening the model for software engineering and long-running agentic work. The company describes more careful planning, sustained execution on multi-step tasks, more reliable behavior inside large codebases, and improved code review and debugging that helps the model catch its own mistakes. Anthropic also highlights the model's ability to handle everyday professional workflows such as financial analyses, research, and working with documents, spreadsheets, and presentations, with autonomous multitasking available through its Cowork research preview.
On evaluations, Anthropic reports that Claude Opus 4.6 sets a new high on the Terminal-Bench 2.0 agentic coding benchmark and leads other frontier models on Humanity's Last Exam, a multidisciplinary reasoning test. On GDPval-AA, which measures economically valuable knowledge work, Anthropic claims Opus 4.6 outperforms the next-best competitor by roughly 144 Elo points and its own predecessor by 190 points, and it also tops BrowseComp for locating hard-to-find information. A standout architectural change is the introduction of a one-the cataloged API limit for an Opus-class model, offered in beta, giving the system room to reason over very large inputs such as full codebases or lengthy document collections.
@ai-sdk/anthropicTransparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Azure Cognitive Services
Delegate complex tasks end-to-end and execute reliably with Claude Opus 4.6, now available in Microsoft Foundry.
Azure Cognitive Services
Claude Opus 4.7 is now available on Snowflake Cortex AI, bringing advanced reasoning and agentic capabilities with Snowflake’s secure governed AI platform, Claude Opus 4.7 is now available on Snowflake Cortex AI, bringing advanced reasoning and agentic capabilities with Snowflake’s secure governed AI platform
Azure Cognitive Services
The release arrives amid weeks of user complaints that Opus 4.6 had quietly gotten worse.
Azure Cognitive Services
Discover more about what's new at AWS with Claude Opus 4.7 is now available in Amazon Bedrock
Pioneer
Claude Opus 4.6 scores 64.3/100 on the benchlm.ai composite capability assessment (31st of 196 ranked models), with pricing at $5 input / $25 output per million tokens and a 1M token context window. The model achieves 39 tokens per second output speed with a 2.23-second first-token latency. It has been superseded by ne Category-level performance shows Agentic at 44.2 (rank 40 of 105, 63rd percentile), Coding at 51.4 (rank 37 of 135, 73rd percentile), Multimodal at 59.6 (rank 30 of 50, 41st percentile), and Knowledge at 57.3 (rank 36 of 160, 78th percentile). Agentic is its lowest-scoring eligible category at 40. The data is current a
AIHubMix
A research summary on Emergent Mind, updated March 12, 2026, evaluates Claude Opus 4.6 as a frontier LLM for enterprise and security applications using rigorous benchmark criteria from the Corecraft and TEE-RedBench studies. The Corecraft-Expanded environment within EnterpriseGym reports task pass rates of 22.10% for C The analysis notes that Claude Opus 4.6 demonstrates advanced reasoning and tool-use capabilities but is constrained by complex multi-step demands in novel enterprise and security workflows, with notable failure modes including poor search strategies, incomplete tool exploration, and challenges in multi-step planning.
AIHubMix
A preprint published on Preprints.org on February 6, 2026 presents a comprehensive comparative analysis of Claude Opus 4.6 alongside OpenAI's GPT-5 series, Google's Gemini 2.5/3 Pro, and GLM-4.6. The paper highlights Claude Opus 4.6's introduction of a 1 million token context window as a significant milestone and ident Performance evaluation spans automated coding, medical informatics, regulatory document processing, and general reasoning benchmarks, positioning Claude Opus 4.6 as a leading model in the early-2026 LLM landscape. The preprint is authored by Satyadhar Joshi and is not peer-reviewed. The paper explicitly names Claude Op