302.AI
SonarSource's independent evaluation explicitly names "Claude Opus 5 Thinking (adaptive thinking mode)" and benchmarks it against Opus 4.8 Thinking on a Java task suite of 4,441 problems spanning HumanEval, MBPP, and ComplexCodeEval. The TLDR figures: an 88.6% pass rate across the 544 HumanEval and MBPP executable-test The review also documents a volume shift: Opus 5 generates 2.3 times as much code as Opus 4.8, with total findings rising 2.7 times — Sonar frames this as changing what "verification" looks like in practice, since the self-verification Anthropic highlights pays off in higher correctness despite larger output. The piece