Azure Cognitive Services
A Vellum analysis of Anthropic's Claude Opus 4.5 system card reports the model scored 80.9% on SWE-bench Verified — surpassing GPT-5.1 and Gemini 3 Pro on real-world GitHub software issue resolution — and 37.6% on ARC-AGI-2, more than doubling GPT 5.1's abstract-reasoning result and beating Gemini 3 Pro by 6%. Vellum a For developers evaluating the exact Claude Opus 4.5 SKU relevant to Azure-hosted deployments, BenchLM's profile (data as of September 4, 2026) records an Anthropic-reported release in November 2025 with a 200K-token context window, $5 input and $25 output per million tokens, throughput of 47 tok/s, and a 1.36-second fi