Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
302.AI logo

Model details

GLM-4.6V

GLM-4.6V is a flagship vision-language model built on a Mixture-of-Experts architecture with 106 billion total parameters and roughly 12 billion activated parameters during inference. This sparse design allows the model to handle complex visual reasoning tasks without activating the full network for every token, making large-scale multimodal reasoning more compute-efficient. The model natively unifies text and image modeling through a dedicated vision encoder that processes multi-resolution inputs up to 4K resolution, enabling it to parse detailed page layouts, charts, and mixed-media documents directly as visual inputs rather than relying solely on extracted text. One of its defining features is native multimodal function calling, which bridges the gap between visual perception and executable downstream actions—a capability specifically designed for building autonomous agents that can see a user interface and take steps within real business workflows. The model also supports interleaved image-text generation and UI reconstruction workflows such as screenshot-to-HTML synthesis, making it practical for product teams working on visual automation tasks.

The model traces its lineage to the GLM-V series introduced in the paper "GLM-4.1V-Thinking and GLM-4.5V: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning," continuing a multimodal training system that leveraged reinforcement learning to scale reasoning capabilities across visual domains. GLM-4.6V achieves state-of-the-art performance on major multimodal benchmarks among models of comparable parameter scales, validating its architectural choices around sparse activation and native tool integration. The series is released under an MIT open-source license with weights publicly available on Hugging Face, lowering the barrier for teams to study, fine-tune, or deploy the model in private environments. Beyond the 106B flagship variant designed for cloud and high-performance cluster scenarios, the family includes a 9B Flash version tailored for local deployment and latency-sensitive applications, offering a practical tiered option for teams with varying infrastructure constraints. These characteristics position GLM-4.6V as a foundation for organizations building vision-capable agents that need both the depth of a large model and the flexibility of open-source deployment.

302.AIglm-4.6vglm

Quick Info

Powered by
Provider
302.AI
Model key
glm-4.6v
Release date
Dec 8, 2025
Last updated
Dec 8, 2025
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.145
Output token cost
$0.43

Limits

Output tokens
32,768 tokens
Context window
128,000 tokens

Transparent token rates

Compare GLM-4.6V pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-4.6V

No articles yet. Fetch the latest news to show it here.

Videos about GLM-4.6V

More models around GLM-4.6V