Currently listed through these providers:
Model details
GLM 4.6V
The GLM-4.6V series is a collection of multimodal large language models engineered to bridge the gap between visual perception and executable action. By integrating native function calling, the architecture allows for direct interaction with tools such as search, cropping, and chart recognition using images, screenshots, or documents as input. This design intent eliminates the need for complex intermediate conversions, reducing information loss and system overhead. The series includes the 106-billion parameter foundation model for high-performance cloud environments and a 9-billion parameter Flash version optimized for low-latency, local deployment, both supporting a 128,000-token context window.
Built upon the lineage of the GLM-V family, these models leverage scalable reinforcement learning to achieve state-of-the-art performance in visual understanding and reasoning. The training process emphasizes versatility, enabling the models to excel in practical business scenarios ranging from automated frontend development and screenshot-to-webpage conversion to image-based shopping assistance. By providing a unified technical foundation for multimodal agents, the series is well-positioned for future-facing applications that require seamless, multi-round visual interaction and reliable, real-world task execution.
Quick Info
Powered by- Provider
- ZenMux
- Model key
- z-ai/glm-4.6v
- Release date
- Dec 8, 2025
- Last updated
- Dec 8, 2025
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.1456
- Output token cost
- $0.4367
Limits
- Output tokens
- 128,000 tokens
- Context window
- 200,000 tokens
Transparent token rates
Compare GLM 4.6V pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about GLM 4.6V
No articles yet. Fetch the latest news to show it here.