FastRouter
Emergent Mind aggregates independent academic evaluations of Claude Opus 4.1 (claude-opus-4-1), positioning it as a large multimodal vision-language model from Anthropic. Two large-scale benchmarking studies, dated September 29 and September 30, 2025, examine the model in high-stakes professional contexts including exp The aggregated findings report Opus 4.1 achieved a mean diagnostic accuracy of 0.01 (1%) on the RadLE v1 benchmark of 50 expert-level spot-diagnosis medical imaging cases, statistically indistinguishable from chance. Reported comparator means include board-certified radiologists at 0.83, radiology trainees at 0.45, GPT