Sulat.com
AI models
Z.AI logo

Model details

GLM-5V-Turbo

GLM-5V-Turbo is positioned as the first member of the GLM family built as a native multimodal agent, meaning visual perception is woven into the architecture rather than added on as a separate interface. It combines image, video, and text inputs with agent-oriented output behaviors such as tool calling, task decomposition, and GUI interaction, extending the lineage that ran through earlier GLM text and vision releases. According to the team's technical report, multimodal perception is treated as a core part of reasoning, planning, tool use, and execution, and the model was developed jointly with Tsinghua University. The documentation describes it as specializing in visual programming, sitting alongside text-focused siblings in the broader Z.ai lineup.

The model's practical sweet spot is turning what it sees into runnable code and agent actions. Reports highlight its ability to reproduce design mockups as near-pixel-perfect HTML in a single pass, and to engage with screenshots, webpages, and GUI surfaces inside agent frameworks. It is optimized for OpenClaw-style high-capacity agentic engineering workflows and preserves competitive text-only coding capability alongside its multimodal strengths. For teams building agents that need to read screens, parse documents, or generate front-end code from visuals, GLM-5V-Turbo is a focused fit, while pure long-horizon language tasks may be better served by the dedicated text models in the same family.

Z.AIglm-5v-turboglm

Quick Info

Powered by
Provider
Z.AI
Model key
glm-5v-turbo
Release date
Apr 1, 2026
Last updated
Apr 1, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$1.20
Output token cost
$4.00

Limits

Output tokens
131,072 tokens
Context window
200,000 tokens

Transparent token rates

Compare glm pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-5V-Turbo

Z.AI

Coverage

Chinese AI startup Zhipu AI has released GLM-5V-Turbo, a multimodal model that processes images, video, and text and is designed for use in agent workflows.

Z.AI

Coverage

Z.ai Launches GLM-5V-Turbo: A Native Multimodal Vision Coding Model Optimized for OpenClaw and High-Capacity Agentic Engineering Workflows Everywhere

Videos about GLM-5V-Turbo

More models around GLM-5V-Turbo