Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
OpenRouter logo

Model details

GLM-4.5V

GLM-4.5V is a next-generation visual reasoning model built on a mixture-of-experts architecture. With a total of 106 billion parameters and 12 billion active parameters, it is designed to move beyond basic perception to provide deep, comprehensive understanding of visual and textual data. The model is engineered to excel in complex problem-solving scenarios, making it a versatile tool for tasks that require high-level reasoning, such as analyzing surveillance footage, interpreting intricate documents, and navigating software interfaces through intelligent agent operations.

The model is built upon the GLM-4.5-Air foundation and follows the technical lineage established by the GLM-4.1V-Thinking series, which emphasizes scalable reinforcement learning to enhance multimodal intelligence. By supporting both thinking and non-thinking modes, it offers flexibility for developers to balance speed and depth depending on the application. Its practical strengths lie in its ability to perform precise actions within digital environments, such as automating workflows across legacy and modern systems, positioning it as a robust solution for developers building the next generation of multimodal AI agents.

OpenRouterz-ai/glm-4.5vglm

Quick Info

Powered by
Provider
OpenRouter
Model key
z-ai/glm-4.5v
Release date
Aug 11, 2025
Last updated
Aug 11, 2025
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.60
Output token cost
$1.80

Limits

Output tokens
16,384 tokens
Context window
65,536 tokens

Transparent token rates

Compare GLM-4.5V pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-4.5V

OpenRouter

Official sourceBenchmark

OpenRouter's product page for Z.ai's GLM 5.3 describes it as a large-scale reasoning model built for complex software engineering and long-horizon agent tasks, supporting text input and output with a 1M-token context window. The page notes improvements over GLM-5.2 in coding and in the balance between performance and t OpenRouter routes GLM 5.3 requests to multiple providers including Reka AI, AkashML, Wafer, DeepInfra, Morph, Friendli, Makora, Cloudflare, Fireworks, Baseten (fp4 and fp8), GMICloud, Modal, Parasail, NovitaAI, DigitalOcean, Together, Phala, SiliconFlow, Decart, Inceptron, Sail Research, Z.ai, and io.net. Performance v

OpenRouter

CoverageRelease Notes

Z.ai's official developer release notes document a 2026 release cadence led by GLM-5.x models. GLM-5.3-Flash (2026-08-26) is described as a native multimodal model with visual capabilities spanning interfaces, rendering results, and interaction feedback, using an efficient hybrid architecture combining linear and spars Earlier entries in the timeline include GLM-5.2 (2026-06-16) with 1M lossless context for long-horizon tasks, and GLM-5.1 (2026-04-07) designed for autonomous runs of up to 8 hours with planning, execution, iterative refinement, and delivery capabilities. The release log also notes GLM-5.3's emergent cybersecurity capa

OpenRouter

Official sourceBenchmark

OpenRouter lists Z.ai's GLM-5V-Turbo as a native multimodal agent foundation model built for vision-based coding and agent-driven tasks, natively handling image, video, and text inputs. The page describes the model as excelling at long-horizon planning, complex coding, and task execution, enabling a "perceive → plan → The model is hosted on OpenRouter by a single provider, Z.ai, with P50 latency of 5.68 seconds and throughput of 28 tokens per second, plus a 79.07% average cache hit rate and 4.15% tool call error rate. OpenRouter's pricing data shows a weighted average effective input price of $0.5331 per 1M tokens against the $1.20

OpenRouter

CoverageBenchmark

Z.ai's GLM-4.5V is hosted on SiliconFlow with metadata listing it as part of the GLM-V family, based on ZhipuAI's GLM-4.5-Air foundation model. According to SiliconFlow's model page, GLM-4.5V achieves state-of-the-art performance on image, video, and document understanding as well as GUI agent operations, and is design SiliconFlow's listing specifies that GLM-4.5V uses a Mixture-of-Experts Transformer architecture with 106B total parameters and 12B activated parameters, FP8 precision, calibrated status, a 66K context length, and no reasoning mode. Critically, SiliconFlow marks the model's state as "Deprecated," indicating it is no lo

Videos about GLM-4.5V

More models around GLM-4.5V