Currently listed through these providers:
Model details
GLM-4.6V
GLM-4.6V sits inside Z.AI's GLM family as a vision-language model aimed at tasks where text reasoning benefits from seeing images. Z.AI has framed the broader 4.6V line as multimodal, pairing language understanding with visual input so the same model can read documents, interpret charts, or describe scenes without routing to a separate vision system. That intent is reinforced by an ecosystem around the release, including a lightweight sibling variant positioned for faster, lower-cost inference when full-scale quality is not required. For developers, this makes GLM-4.6V a practical choice for assistants and pipelines that need to mix natural language with visual context in a single API call rather than stitching together specialized models. In practical terms, GLM-4.6V is shaped for production multimodal workflows: image-conditioned question answering, visual document parsing, UI and screenshot understanding, and reasoning chains that reference what is shown in an image. Because the model is part of the open-weights GLM lineage, it can be self-hosted, fine-tuned, or deployed behind private infrastructure for teams with data-residency or compliance needs. It fits well alongside tool-using agents that need to look at a picture, read a chart, or check a form before deciding the next action, giving builders a single model that handles both the reading and the reasoning step.
Within the GLM-4.6V release wave, Z.AI has signaled continued investment in compact multimodal variants, with a smaller Flash sibling explicitly designed to preserve the family's multimodal behavior at lower compute cost for latency-sensitive applications. That product structure suggests a roadmap in which heavier GLM-4.6V deployments handle the deepest multimodal reasoning, while lighter variants cover high-volume or real-time use cases under the same API surface. For teams evaluating the lineup, the practical pattern is to prototype with the full GLM-4.6V for quality-sensitive workflows and reach for lighter siblings when throughput, cost per request, or response time dominate the requirement. Together, these models extend the GLM family's reach into a broader range of production scenarios where understanding both text and images in the same context window matters.
Quick Info
Powered by- Provider
- OpenRouter
- Model key
- z-ai/glm-4.6v
- Release date
- Dec 8, 2025
- Last updated
- Dec 8, 2025
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.30
- Output token cost
- $0.90
Limits
- Output tokens
- 32,768 tokens
- Context window
- 131,072 tokens