OpenRouter
The LLM-Stats aggregator page for Qwen2.5 VL 72B Instruct places the model at composite rank 249 overall, with capability-tier standings of 78/119 in Long Context, 130/208 in Vision, 147/244 in Healthcare, 240/362 in Reasoning, and 263/327 in Math. The page also tracks how the model holds up as conversation length grow Per-benchmark scores sourced from the model's own scorecard, paper, or official blog posts include DocVQA rank 1 at 0.96, Android Control Low EM rank 1 at 0.94, ChartQA rank 3 at 0.90, and OCRBench rank 10 at 0.89, with 26 additional benchmarks listed. LLM-Stats notes that these scores are self-reported by the model pr