Currently listed through these providers:
Model details
GLM-4.6V
The GLM-4.6V series serves as a foundational architecture for multimodal reasoning, designed to bridge the gap between visual perception and actionable output. By integrating native function calling, the model allows for direct interaction with external tools—such as search, cropping, or chart recognition—using images and documents as primary inputs. This design intent focuses on reducing the complexity and potential information loss associated with traditional text-based conversions, enabling the model to perform tasks like frontend replication, UI mockup generation, and complex visual interaction development with greater precision.
Built upon a lineage that emphasizes scalable reinforcement learning and versatile multimodal reasoning, the series offers two distinct versions to suit different deployment needs. The 106B foundation model is engineered for high-performance cloud clusters, while the 9B Flash variant provides a lightweight alternative for local, low-latency applications. These models demonstrate state-of-the-art performance in visual understanding and reasoning across various benchmarks, providing a robust technical foundation for developers building agents that require autonomous planning and execution in real-world business environments.
Quick Info
Powered by- Provider
- Zhipu AI
- Model key
- glm-4.6v
- Release date
- Dec 8, 2025
- Last updated
- Dec 8, 2025
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.30
- Output token cost
- $0.90
Limits
- Output tokens
- 32,768 tokens
- Context window
- 128,000 tokens
Transparent token rates
Compare glm pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about GLM-4.6V
No articles yet. Fetch the latest news to show it here.