Currently listed through these providers:
Model details
GLM-5V-Turbo
GLM-5V-Turbo is positioned by its developer as a first-of-its-kind multimodal coding foundation model, purpose-built for vision-based coding tasks rather than general conversation. It natively accepts images, video, text, and files as input and produces text output, letting a single model interpret design mockups, reference screenshots, screen recordings, or mixed file inputs and translate them directly into code. The model is deeply optimized for agent workflows, integrating with coding agents such as Claude Code and OpenClaw to close the loop of understanding the visual environment, planning a sequence of actions, and executing tasks end to end.
In practice, GLM-5V-Turbo is aimed at builders who want a long-context, vision-aware collaborator for agentic coding pipelines. It offers multiple thinking modes so teams can trade off depth of reasoning against latency, and pairs that with vision comprehension, streaming output, function calling, and context caching to keep multi-step runs efficient. With a 200K context window and a 128K maximum output, it can sustain long-horizon planning across complex codebases while staying responsive through tool calls and streamed responses, making it a strong fit for frontend recreation from designs, repository-scale refactors driven by visual references, and orchestrated agent loops where vision, reasoning, and tool use all need to coexist.
Quick Info
Powered by- Provider
- OpenRouter
- Model key
- z-ai/glm-5v-turbo
- Release date
- Apr 1, 2026
- Last updated
- Apr 1, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $1.20
- Output token cost
- $4.00
Limits
- Output tokens
- 131,072 tokens
- Context window
- 202,752 tokens