Kilo Gateway
Chinese AI startup Zhipu AI has released GLM-5V-Turbo, a multimodal model that processes images, video, and text and is designed for use in agent workflows.
Model details
GLM-5V-Turbo is a vision-coding base model from Zhipu AI, the Chinese AI company also known internationally as Z.ai, and is positioned as the first multimodal entry in the GLM coding line. It is designed to read visual design inputs, such as interface mockups and screenshots, and translate them directly into working front-end code in a single pass, unifying perception and software generation rather than chaining a vision model to a separate code model. Independent coverage describes the system as a "vision coding agent" aimed squarely at developers and design-to-code workflows, distinguishing it from text-only coding models by treating the screen itself as the prompt.
The model carries forward lineage from earlier GLM-5 releases while specializing that foundation for GUI understanding, and early reporting highlights strong results in coding and GUI-agent benchmarks, reflecting a focus on tasks where a model must both interpret a visual interface and produce reliable, executable output. Practical fit centers on front-end prototyping, converting mockups into starter projects, and powering agents that act on graphical user interfaces, where combining vision and code in one model reduces the latency and error of multi-stage pipelines. The result is a coding model whose differentiator is not raw language ability but tight, end-to-end integration of what it sees and what it ships.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Kilo Gateway
Chinese AI startup Zhipu AI has released GLM-5V-Turbo, a multimodal model that processes images, video, and text and is designed for use in agent workflows.
Kilo Gateway
Z.ai Launches GLM-5V-Turbo: A Native Multimodal Vision Coding Model Optimized for OpenClaw and High-Capacity Agentic Engineering Workflows Everywhere
OpenRouter
The alphaXiv technical report "GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents" (abs/2604.26752v1) provides a model-specific technical reference independent of marketing coverage. Its abstract frames the work as a step toward native foundation models for multimodal agents, arguing that agentic capa The report states that GLM-5V-Turbo integrates multimodal perception directly into reasoning, planning, tool use, and execution rather than treating it as an auxiliary interface to a language model. It documents improvements across model design, multimodal training, reinforcement learning, toolchain expansion, and inte
TensorX
The Opper AI Z.ai release tracker explicitly lists GLM-5V-Turbo with an April 1, 2026 release date, a 205K context window, an intelligence score of 24, and per-million-token pricing of $1.20 for input and $4.00 for output. The tracker cross-checks release dates and intelligence scores against vendor announcements via A The same tracker situates GLM-5V-Turbo in Z.ai's broader 2026 release cadence, appearing between the March 15, 2026 GLM-5-Turbo launch and the April 7, 2026 GLM-5.1 release, and one of fourteen Z.ai models indexed on the platform. This positioning is useful for developers tracking the GLM family roadmap, as it shows GL