Z.AI
Z.ai officially introduced and open-sourced the GLM-4.6V series on December 8, 2025, with GLM-4.6V-Flash (9B) positioned as a lightweight variant optimized for local deployment and low-latency applications. The series scales its context window to 128k tokens during training and achieves state-of-the-art visual understanding and reasoning among models of similar parameter scales. Model weights and code are released on Hugging Face and GitHub under MIT license. The defining innovation highlighted in the Z.ai blog is native multimodal function calling, allowing images, screenshots, and document pages to pass directly as tool parameters without intermediate text conversion, and letting the model visually interpret returned outputs like charts and web snapshots. GLM-4.6V-Flash (9B) reportedly outperforms Qwen3-VL-8B at comparable parameter scales, enabling agent use cases such as mixed text-image content creation, visual web search, and complex document understanding. The model accepts video, image, text, and file inputs and produces text output.