LLM Gateway
Compare GLM-4.5V vs Kimi K2.5: input $0.55/M vs $0.6/M, output $2.19/M vs $2.5/M tokens. GLM-4.5V is 12% cheaper overall. Full API cost breakdown, context window, and benchmark comparison.
Model details
GLM-4.5V is a vision-language model built on a mixture-of-experts architecture, designed to handle complex reasoning tasks that go beyond basic perception. With a total of 106 billion parameters and 12 billion active parameters, the model is engineered to provide high-level intelligence for multimodal applications. It is specifically optimized for deep analysis of images, videos, and intricate documents, making it a versatile tool for developers who need to extract insights from diverse visual data or automate interactions across web, desktop, and mobile interfaces.
The model is part of the GLM-V family and draws its technical foundation from the GLM-4.5-Air text model, continuing the development path established by the GLM-4.1V-Thinking series. By utilizing scalable reinforcement learning, the model achieves state-of-the-art performance among open-source vision-language models of its scale. Its design supports both thinking and non-thinking modes, allowing it to function effectively as a multimodal agent capable of navigating legacy systems or generating detailed, timestamped reports from long-form surveillance footage.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
LLM Gateway
Compare GLM-4.5V vs Kimi K2.5: input $0.55/M vs $0.6/M, output $2.19/M vs $2.5/M tokens. GLM-4.5V is 12% cheaper overall. Full API cost breakdown, context window, and benchmark comparison.