DevPass (LLM Gateway)
Qwen3.8-27B scores 52 on the Artificial Analysis Intelligence Index at maximum reasoning effort and 61.7 on SWE-bench Pro per Qwen's own evaluation. The model leads its comparison set on agentic coding and computer-use benchmarks while trailing frontier models on Humanity's Last Exam and GPQA Diamond, reflecting a prof The Qwen3.8-27B vs Qwen3.6-27B vs Qwen3.7-Plus tables show DeepSWE 1.1 moving from 13.3 to 42.2, QwenSWEBench from 49.3 to 79.0, OSWorld-Verified from 63.9 to 84.3, Browser-use WebArena-Verified from 48.8 to 64.8, Mobile-use AndroidWorld from 70.3 to 81.9, RecreationBench from 29.8 to 47.1, SWE-MM from 25.7 to 38.6, an