Z.AI
Compare GLM-4.5V vs Kimi K2.5: input $0.55/M vs $0.6/M, output $2.19/M vs $2.5/M tokens. GLM-4.5V is 12% cheaper overall. Full API cost breakdown, context window, and benchmark comparison.
Model details
GLM-4.5V is a vision-language foundation model built on a Mixture-of-Experts architecture with 106 billion total parameters and 12 billion activated per token, making it highly efficient for diverse multimodal tasks. Based on the GLM-4.5-Air text foundation model, it continues the technical lineage of GLM-4.1V-Thinking and achieves state-of-the-art performance among models of the same scale across 42 public vision-language benchmarks. The model is engineered for multimodal agent applications, covering image understanding, video comprehension, document parsing, OCR, front-end web coding, grounding, and spatial reasoning. A distinguishing capability is its hybrid inference mode, which offers both a "thinking mode" for deep reasoning and a "non-thinking mode" for fast responses, with the reasoning behavior toggleable via a parameter.
The development approach draws on scalable reinforcement learning to advance versatile multimodal reasoning, as explored in the GLM-V Team's research. This training methodology enables the model to handle complex real-world scenarios, including autonomous interaction with graphical user interfaces through intelligent agents. The model ships with open weights on Hugging Face, allowing developers to run it locally or deploy it across various platforms. Its design emphasizes practical applications—extracting insights from complex documents, interpreting images and videos accurately, and serving as a backbone for sophisticated AI agents that need to perceive and act within multimodal environments.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Z.AI
Compare GLM-4.5V vs Kimi K2.5: input $0.55/M vs $0.6/M, output $2.19/M vs $2.5/M tokens. GLM-4.5V is 12% cheaper overall. Full API cost breakdown, context window, and benchmark comparison.