Currently listed through these providers:
Model details
GLM-4.6V-Flash
GLM-4.6V-Flash is the lightweight 9-billion-parameter variant of Zhipu AI's GLM-4.6V multimodal series, positioned alongside the larger 106B foundation model for cloud and high-performance deployments. As the entry-level member of the family, Flash is explicitly optimized for local deployment and low-latency applications, making it a practical option for developers who want capable visual reasoning without relying on remote infrastructure. The model handles text, image, and video inputs and produces text outputs, fitting the typical pattern of a vision-language assistant aimed at chat, document understanding, and visual question answering scenarios.
A defining feature of the GLM-4.6V series, and therefore of Flash as well, is the integration of native multimodal Function Calling, which allows images, screenshots, and document pages to be passed directly as tool parameters rather than being converted through intermediate text steps. This design is intended to bridge visual perception and executable action, giving the model a unified foundation for multimodal agent workflows. The creator's release post frames the family as state-of-the-art among comparably scaled models on visual understanding and reasoning, and notes the cataloged API limit training context that supports long documents and extended multimodal sessions. Being open-sourced on Hugging Face and GitHub, GLM-4.6V-Flash is well suited for teams building on-device assistants, tool-augmented agents, or research projects that benefit from a compact yet multimodal-capable base model.
Quick Info
Powered by- Provider
- Zhipu AI
- Model key
- glm-4.6v-flash
- Release date
- Dec 8, 2025
- Last updated
- Dec 8, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 32,768 tokens
- Context window
- 128,000 tokens
Latest news about GLM-4.6V-Flash
No articles yet. Fetch the latest news to show it here.