Sulat.com
AI models
Zhipu AI logo

Model details

GLM-4.6V

The GLM-4.6V series serves as a foundational architecture for multimodal reasoning, designed to bridge the gap between visual perception and actionable output. By integrating native function calling, the model allows for direct interaction with external tools—such as search, cropping, or chart recognition—using images and documents as primary inputs. This design intent focuses on reducing the complexity and potential information loss associated with traditional text-based conversions, enabling the model to perform tasks like frontend replication, UI mockup generation, and complex visual interaction development with greater precision.

Built upon a lineage that emphasizes scalable reinforcement learning and versatile multimodal reasoning, the series offers two distinct versions to suit different deployment needs. The 106B foundation model is engineered for high-performance cloud clusters, while the 9B Flash variant provides a lightweight alternative for local, low-latency applications. These models demonstrate state-of-the-art performance in visual understanding and reasoning across various benchmarks, providing a robust technical foundation for developers building agents that require autonomous planning and execution in real-world business environments.

Zhipu AIglm-4.6vglm

Quick Info

Powered by
Provider
Zhipu AI
Model key
glm-4.6v
Release date
Dec 8, 2025
Last updated
Dec 8, 2025
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.30
Output token cost
$0.90

Limits

Output tokens
32,768 tokens
Context window
128,000 tokens

Transparent token rates

Compare glm pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-4.6V

No articles yet. Fetch the latest news to show it here.

Videos about GLM-4.6V

More models around GLM-4.6V