Currently listed through these providers:
Model details
GLM 5 Vision Turbo
GLM 5 Vision Turbo is Zhipu's first native multimodal foundation model aimed specifically at visual programming and agent-driven coding tasks. It accepts images, video, and text as native inputs and produces text output, positioning it as a tool for systems that need to interpret visual environments and translate them into action. The model is described as deeply adapted to agent workflows, designed to collaborate closely with frameworks such as Claude Code and OpenClaw to complete a closed loop of understanding the environment, planning actions, and executing tasks. Its design focus on long-horizon planning and complex programming makes it a strong fit for end-to-end automation rather than single-turn visual question answering.
Practically, GLM 5 Vision Turbo is well suited to developers who want to feed design mockups, UI screenshots, or video walkthroughs straight into a coding pipeline and have the model generate working front-end code or orchestrate subsequent tool calls. Its substantial context window allows it to hold long working memory across multi-step agent sessions, and its speed profile is competitive for interactive developer tooling. A real-world demonstration highlighted by independent coverage showed the model converting design mockups directly into executable front-end code, illustrating its core value proposition of bridging visual design and production software without manual translation steps.
Quick Info
Powered by- Provider
- AIHubMix
- Model key
- glm-5v-turbo
- Release date
- May 9, 2026
- Last updated
- May 9, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.7042
- Output token cost
- $3.09848
Limits
- Output tokens
- 128,000 tokens
- Context window
- 200,000 tokens