Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Z.AI logo

Model details

GLM-4.6V-Flash

GLM-4.6V-Flash sits at the lightweight end of the GLM-4.6V multimodal family, a series designed to combine visual perception with executable action through first-party Function Calling. Z.AI positions the 9B-parameter variant for local deployment and low-latency applications, in contrast to the larger 106B foundation model aimed at cloud and high-performance clusters. The series was trained with an extended context window suited to long documents, image sequences, and video, and it is the first GLM iteration to ship native multimodal tool calling rather than relying on text-only intermediate steps for invoking functions on visual inputs.

Because the weights are released openly, GLM-4.6V-Flash is practical to self-host and fine-tune for agentic scenarios where a model needs to perceive a screen, document, or scene and then act on it. Its native function calling lets images, screenshots, and document pages be passed directly as tool parameters, reducing the information loss that comes from routing visual content through separate text-only stages. For teams building local assistants, GUI agents, or document-processing pipelines that require both visual understanding and reliable tool execution without the cost of a frontier-scale model, this lightweight variant offers a balanced fit between capability and deployable footprint.

Z.AIglm-4.6v-flashglm-flash

Quick Info

Powered by
Provider
Z.AI
Model key
glm-4.6v-flash
Release date
Dec 8, 2025
Last updated
Dec 8, 2025
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
32,768 tokens
Context window
128,000 tokens

Latest news about GLM-4.6V-Flash

Z.AI

Official sourceAnnouncement

Z.ai officially introduced and open-sourced the GLM-4.6V series on December 8, 2025, with GLM-4.6V-Flash (9B) positioned as a lightweight variant optimized for local deployment and low-latency applications. The series scales its context window to 128k tokens during training and achieves state-of-the-art visual understanding and reasoning among models of similar parameter scales. Model weights and code are released on Hugging Face and GitHub under MIT license. The defining innovation highlighted in the Z.ai blog is native multimodal function calling, allowing images, screenshots, and document pages to pass directly as tool parameters without intermediate text conversion, and letting the model visually interpret returned outputs like charts and web snapshots. GLM-4.6V-Flash (9B) reportedly outperforms Qwen3-VL-8B at comparable parameter scales, enabling agent use cases such as mixed text-image content creation, visual web search, and complex document understanding. The model accepts video, image, text, and file inputs and produces text output.

Z.AI

CoverageRelease Notes

Zhipu AI open-sourced the GLM-4.6V series, with GLM-4.6V-Flash positioned as a 9B parameter variant tuned for local deployment and low-latency use. The series treats visual inputs—images, video, and document pages—as first-class inputs for agents rather than text add-ons. This Flash variant is designed to run efficiently on local hardware while supporting the full multimodal tool-use capabilities of the family. GLM-4.6V-Flash shares its family's 128K token training context window, enabling roughly 150 pages of dense documents, 200 slide pages, or one hour of video in a single pass. Its key technical contribution is native multimodal Function Calling, where images, screenshots, and document pages pass directly as tool parameters rather than being converted to text descriptions. Zhipu AI extends the Model Context Protocol with URL-based multimodal handling to support this, allowing tools to return search result grids, charts, and rendered web pages that the model fuses with text in the same reasoning chain.

Z.AI

Coverage

VentureBeat reported on December 8, 2025, that Z.ai (Zhipu AI) released the GLM-4.6V series of open-source vision-language models, including the 9-billion-parameter GLM-4.6V-Flash aimed at low-latency, local applications. The series introduces native function calling in a vision-language model, enabling direct tool use with visual inputs such as search, cropping, and chart recognition across a 128,000-token context window. The models achieve state-of-the-art results across more than 20 benchmarks, and both GLM-4.6V and GLM-4.6V-Flash ship under the permissive MIT license for free commercial and non-commercial use, including air-gapped deployments. Weights and documentation are hosted on Hugging Face, with supporting code on GitHub, and access is available via an OpenAI-compatible API, a Z.ai web demo, and a Hugging Face Spaces desktop assistant app.

Z.AI

CoverageRelease Notes

Pandaily reported on December 8, 2025, that Zhipu AI released and open-sourced the GLM-4.6V multimodal model series, with GLM-4.6V-Flash (9B) featured as a lightweight variant optimized for on-device deployment and low-latency use cases. The series is built around an "images as parameters, results as context" philosophy, introducing native multimodal toolcalling that removes the need to convert visual inputs to text before invoking tools. Across more than 30 multimodal benchmarks including MMBench, MathVista, and OCRBench, GLM-4.6V shows substantial gains over its predecessor, with the compact 9B Flash variant outperforming Qwen3-VL-8B and the 106B-A12B flagship rivaling Qwen3-VL-235B at roughly half the parameters. The series targets agent workflows spanning mixed-format content generation, product recognition, price-value recommendations, and complex document understanding.

Z.AI

Official sourceDocumentation

The official Z.AI developer documentation enumerates GLM-4.6V-Flash as a distinct variant of the GLM-4.6V series, alongside the flagship GLM-4.6V and GLM-4.6V-FlashX. Flash is positioned as "Lightweight, Completely Free" with 128K context length, video/image/text/file input, and text output. The page confirms native multimodal function calling that bridges visual perception and executable action for multimodal agents. The documentation details capabilities including intelligent image-text content creation and layout, visual web search and rich media report generation, frontend replication and visual interaction, and long-context understanding. GLM-4.6V-Flash accepts multimodal inputs such as papers, reports, and slides, and can pass visual artifacts directly as tool parameters while visually interpreting tool-returned results like charts and rendered screenshots. This closes the perception-to-execution loop for agent workflows.

Videos about GLM-4.6V-Flash

More models around GLM-4.6V-Flash