Currently listed through these providers:
Model details
GLM-4.6V
The GLM-4.6V series represents a significant evolution in multimodal design, utilizing a Mixture-of-Experts architecture to balance high-level performance with operational efficiency. The flagship 106B model is engineered for demanding cloud and high-performance cluster environments, while the 9B Flash version provides a streamlined alternative for local deployment and low-latency tasks. By integrating native function calling, the series bridges the gap between visual perception and executable action, allowing the models to process images, screenshots, and documents directly as tool parameters. This design intent focuses on creating a unified foundation for multimodal agents capable of complex interactions, such as converting visual layouts into web code or performing image-based product searches.
Building upon the technical lineage of the GLM-V series, these models benefit from a training system that emphasizes deep visual understanding and reasoning. The architecture supports multi-resolution image processing up to 4K, ensuring precision in tasks that require high-fidelity visual analysis. By moving away from traditional text-only tool use, the models minimize information loss and system complexity, enabling more fluid, multi-round visual interactions. This technical approach positions the series as a robust choice for developers building intelligent agents that require both sophisticated reasoning and the ability to interact with external digital tools in real-world business scenarios.
Quick Info
Powered by- Provider
- DevPass (LLM Gateway)
- Model key
- glm-4.6v
- Release date
- Dec 8, 2025
- Last updated
- Dec 8, 2025
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.30
- Output token cost
- $0.90
Limits
- Output tokens
- 32,768 tokens
- Context window
- 131,072 tokens
Transparent token rates
Compare GLM-4.6V pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.