Z.AI
Z.AI released GLM-5V-Turbo as its first multimodal coding foundation model, built for vision-based coding and agentic tasks with native support for image, video, and text inputs. The model targets long-horizon planning, complex coding, and full perceive-plan-execute agent loops, per the Agent Native analysis. GLM-5V-Turbo was trained with a fully fused text+vision pipeline from pretraining through fine-tuning, using a CogViT visual encoder for image and video understanding. The analysis reports a BridgeBench SpeedBench score of 5 at 221.2 tokens/sec and notes synergy with Claude Code and OpenClaw, positioning it for GUI agents and autonomous UI exploration workflows.