Z.AI
Z.AI (formerly Zhipu) announced GLM-4.5V on August 13, 2025, as an open-source vision-language model engineered for multimodal reasoning across images, video, long documents, charts, and GUI screens. The model uses a 106B total / 12B active MoE architecture designed to pair high accuracy with practical latency and deployment cost. It follows the earlier GLM-4.1V-9B-Thinking release and scales that approach to enterprise workloads while keeping developer ergonomics central. Built on the GLM-4.5-Air text base and extending the GLM-4.1V-Thinking lineage, GLM-4.5V delivers SOTA performance among similarly sized open-source VLMs across 41 public multimodal evaluations. Access channels include Hugging Face, GitHub, the Z.AI API platform, and Z.AI Chat, ensuring broad developer reach. Capabilities span image reasoning with localization, video shot segmentation and event recognition, GUI tasks including screen reading and icon detection, complex chart and long-document analysis, and precise spatial grounding with bounding coordinates.