Zhipu AI
Chinese AI startup Zhipu AI has released GLM-5V-Turbo, a multimodal model that processes images, video, and text and is designed for use in agent workflows.
Model details
GLM-5V-Turbo represents a deliberate architectural shift for the GLM family. Where earlier models like GLM-4V handled images as an extension and GLM-5 sharpened text-based reasoning, GLM-5V-Turbo treats vision as a foundational layer woven into the model's design from the start rather than a later addition. A proprietary vision encoder enables it to process design mockups, video, images, and text within a unified framework, and the model is built specifically for agentic workflows that move beyond simple input-output responses toward integrated perception, planning, and execution. For developers, this architecture means the model can receive a screenshot of a UI design and generate functional front-end code directly, handling the translation between visual intent and executable output in a single pipeline.
The model's lineage traces through the rapid GLM iteration cycle, which saw the family progress from GLM-4.5 through GLM-5 within less than a year before introducing this multimodal variant. Its optimization for agent frameworks like OpenClaw and Claude Code positions it as a foundation for complex, long-horizon programming tasks where a model needs to maintain context across extended interactions and decompose multi-step problems. GLM-5V-Turbo demonstrates strong results on both multimodal coding benchmarks and GUI agent evaluation suites while retaining capabilities in pure text-based coding, making it suitable for teams building automated development workflows that combine visual understanding with sophisticated code generation and tool interaction.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Zhipu AI
Chinese AI startup Zhipu AI has released GLM-5V-Turbo, a multimodal model that processes images, video, and text and is designed for use in agent workflows.
Zhipu AI
Z.ai Launches GLM-5V-Turbo: A Native Multimodal Vision Coding Model Optimized for OpenClaw and High-Capacity Agentic Engineering Workflows Everywhere
Zhipu AI
GLM-5V-Turbo is Z.AI's first multimodal coding foundation model, built for vision-based coding tasks. It natively processes multimodal inputs including images, video, text, and files, while excelling at long-horizon planning, complex coding, and action execution. Deeply optimized for agent workflows, it works seamlessl