Baseten
Baseten published a technical write-up on July 29, 2026 detailing how it post-trained vision capabilities onto the otherwise text-only GLM 5.2 open model. Rather than training a full vision tower, the team trained only a two-layer MLP projector of roughly 50 million parameters on top of the existing vision encoder from The resulting model reportedly reaches MMMU-Pro parity with Claude 4.5 Haiku at 55%, and the team notes an emergent generalization effect: GLM 5.2 could identify people who were never explicitly labeled in their training data. The post positions the recipe as a lightweight, reproducible path for adding multimodal input