SiliconFlow (China)
Reproducible benchmark results for Qwen/Qwen3-8B were evaluated across 13 tasks on a 16.4 GB safetensors checkpoint, providing independent verification of the model's capabilities. The strongest showing was GSM8K at 84.9% exact match (rank 10 of 28), followed by IFEval at 81.7% prompt-level strict accuracy and MMLU at 73.0%. These measurements give developers third-party validation of Qwen3-8B's math, instruction-following, and general knowledge performance. Additional reproducible scores include MBPP at 65.6% pass@1, MMLU-Pro at 57.7%, MGSM at 47.9%, and EQ-Bench at 75.8, while GPQA Diamond reached 27.8% and the MBPP (Instruct) variant scored 0.0%. The model ranked in the top three on none of the 13 benchmarks evaluated, with its weakest results on instruction-formatted MBPP, giving a clear picture of where Qwen3-8B excels versus where it struggles in real coding workloads.