NovitaAI
The MLB clinical benchmark evaluates Baichuan-M2-32B alongside 10 leading models across medical knowledge, safety and ethics, medical-record understanding, smart services, and smart healthcare. The benchmark covers 64 clinical specialties using 22 datasets, 17 of them newly curated, and a curation process involving 300 Baichuan-M2-32B achieved a 90.6% score in the benchmark’s Safety and Ethics dimension, which the paper describes as exceptional given the model’s comparatively small size. The result supports a focused technical conclusion about targeted training and medical-safety performance, but it is third-party benchmark evidence