Kilo Gateway
A peer-reviewed study published January 7, 2026 in Frontiers in Medicine evaluated OpenAI o4-mini-high alongside Claude 4 Opus, Gemini 2.5 Pro, and Qwen 3 on the 200-item New England Journal of Medicine Image Challenge. The variant achieved the highest overall accuracy at 94%, with consistently strong performance acros An error analysis found that 83.3% of o4-mini-high's mistakes reflected lapses in diagnostic logic rather than input processing, and simple prompting techniques including chain-of-thought and few-shot learning corrected over half of these errors. While the evaluation is domain-specific and does not constitute release o