Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Venice AI logo

Model details

GLM 5V Turbo

GLM 5V Turbo is a vision-language foundation model purpose-built for coding tasks, marking Z.ai's first natively multimodal design rather than a vision module bolted onto a text model. It accepts images, video, text, and file inputs and produces text output, combining a CogViT vision encoder with a Multi-Token Prediction architecture on top of the Turbo variant of the GLM-5 line. The design intent is to close the loop between visual perception and code generation: a screenshot, mockup, wireframe, or UI recording is parsed for layout, color, component hierarchy, and interaction logic, then emitted as a runnable front-end project. The model is positioned for long-horizon planning, complex coding, and action execution, and is deeply integrated with agent frameworks such as Claude Code and OpenClaw, where it handles environment understanding, action planning, and tool invocation end to end.

Practical strengths show up most clearly on the Design2Code benchmark, where GLM 5V Turbo reaches 94.8, opening a roughly seventeen-point lead over comparable frontier multimodal models on translating visual designs into HTML and CSS, and reports place it ahead of Claude Opus 4.6 on multimodal evaluation suites. The same family includes the text-only GLM-5 at 744B parameters and the text-optimized GLM-5-Turbo, with GLM 5V Turbo adding native multimodal understanding without changing the underlying API surface or pricing tier. It supports multiple thinking modes, streaming responses, function calling, and intelligent context caching, and ships with a large working memory suited to multi-step agentic sessions. Forward-looking, it is aimed at teams that want an agent-ready vision coder that can ingest screenshots and video, drive tool use, and ship production front-end output in a single pass.

Venice AIz-ai-glm-5v-turboglm

Quick Info

Powered by
Provider
Venice AI
Model key
z-ai-glm-5v-turbo
Release date
Apr 1, 2026
Last updated
Jun 11, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$1.50
Output token cost
$5.00

Limits

Output tokens
32,768 tokens
Context window
200,000 tokens

Transparent token rates

Compare GLM 5V Turbo pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM 5V Turbo

Venice AI

CoverageBenchmark

A CSPaper benchmark snapshot dated 20 July 2026 evaluates GLM-5V-Turbo alongside Gemini 3.1 Pro and GPT-5.6 Terra across 38 public venue and track rows drawn from a corpus of 700 expert-calibrated papers. GLM-5V-Turbo records the lowest NMAE at 0.165, the smallest mean bias at 0.000 (near zero), and an SRC ranking agre The same article notes that Gemini 3.1 Pro is the strongest ranker in the group with the most consistent paper ordering, while GPT-5.6 Terra produces the longest reviews but scores more strictly and ranks lower than GLM-5V-Turbo on accuracy. The piece frames the comparison as a multi-metric snapshot rather than a singl

Venice AI

Coverage

Chinese AI startup Zhipu AI has released GLM-5V-Turbo, a multimodal model that processes images, video, and text and is designed for use in agent workflows.

Venice AI

Coverage

Z.ai Launches GLM-5V-Turbo: A Native Multimodal Vision Coding Model Optimized for OpenClaw and High-Capacity Agentic Engineering Workflows Everywhere

ZenMux

Coverage

The alphaXiv preprint presents GLM-5V-Turbo as a step toward native foundation models for multimodal agents, where agentic capability requires perceiving, interpreting, and acting over heterogeneous contexts such as images, videos, webpages, documents, and GUIs. The abstract frames multimodal perception as integrated i The paper summarizes main improvements across model design, multimodal training, reinforcement learning, toolchain expansion, and integration with agent frameworks, reporting strong performance in multimodal coding, visual tool use, and framework-based agentic tasks while preserving competitive text-only coding capabil

Videos about GLM 5V Turbo

More models around GLM 5V Turbo