Sulat.com
AI models
Zhipu AI Coding Plan logo

Model details

GLM-5V-Turbo

GLM-5V-Turbo represents a collaborative effort between Z.ai and Tsinghua University to build what its technical report describes as a native foundation model for multimodal agents. Rather than treating visual understanding as an auxiliary interface bolted onto a language model, the team integrated multimodal perception as a core component of the model's reasoning, planning, tool use, and execution pipelines. The arXiv paper details improvements across model design, multimodal training, reinforcement learning, and toolchain expansion, with the explicit goal of enabling agents that can perceive, interpret, and act over heterogeneous contexts such as images, videos, webpages, documents, and graphical user interfaces.

The model is positioned for practical, vision-centric coding workflows, with reporting highlighting its ability to translate design mockups directly into executable front-end code. On BridgeBench SpeedBench, it achieved approximately 221.2 tokens per second, landing it among the faster multimodal systems tracked. Independent coverage from Artificial Analysis cited via release trackers places its context window around 205K tokens. GLM-5V-Turbo fits well for teams building coding assistants, GUI-grounded agents, and design-to-code pipelines that need to reason jointly over visual inputs and programmatic generation, while still preserving competitive text-only coding capability.

Zhipu AI Coding Planglm-5v-turboglm

Quick Info

Powered by
Provider
Zhipu AI Coding Plan
Model key
glm-5v-turbo
Release date
Apr 1, 2026
Last updated
Apr 1, 2026
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
131,072 tokens
Context window
200,000 tokens

Latest news about GLM-5V-Turbo

Zhipu AI Coding Plan

CoverageRelease Notes

According to the Opper AI release tracker, GLM-5V-Turbo was released on 1 April 2026 as part of Z.ai's (Zhipu AI) multimodal line, distributed via the Zhipu AI Coding Plan endpoint at open.bigmodel.cn. The tracker's release table lists the model alongside other GLM-5 family variants and is the only candidate source tha The tracker records GLM-5V-Turbo with a 205K-token context window, an intelligence index of 35, and pricing of $1.20 per million input tokens and $4.00 per million output tokens, with the intelligence score and pricing attributed to Artificial Analysis rather than first-party Z.ai sources. No vendor documentation, API

Zhipu AI Coding Plan

Coverage

Chinese AI startup Zhipu AI has released GLM-5V-Turbo, a multimodal model that processes images, video, and text and is designed for use in agent workflows.

Zhipu AI Coding Plan

CoverageBenchmark

GLM-5V-Turbo is Z.AI’s first multimodal coding foundation model, built for vision-based coding tasks. It also hit #5 on BridgeBench SpeedBench with 221.2 tokens/sec, faster than Gemini 3.1 Pro…

Zhipu AI Coding Plan

Official sourceRelease Notes

Follow along with updates across Z.AI’s models

Videos about GLM-5V-Turbo

More models around GLM-5V-Turbo