Currently listed through:
Model details
GLM 4.5V
GLM 4.5V is a vision-language model from the GLM-V Team that extends the GLM-4.1V-Thinking approach into a more versatile multimodal reasoning system, as documented in the arXiv paper "GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning." The architecture is built on ZhipuAI's next-generation GLM-4.5-Air text foundation model and uses a mixture-of-experts design with 106B total parameters and 12B active parameters, letting it handle multimodal inputs efficiently without activating the full parameter set on every token. Its training approach builds on scalable reinforcement learning to strengthen multimodal chain-of-thought reasoning across diverse input types.
In practical terms, GLM 4.5V is positioned for image, video, and long document understanding as well as GUI agent operations that involve interpreting screen interfaces and acting on them. Fireworks AI reports that it achieves state-of-the-art performance among same-scale models on 42 public vision-language benchmarks, suggesting a broad and competitive capability profile rather than narrow specialization. The combination of MoE efficiency, broad multimodal coverage, and GUI-agent readiness makes it a strong fit for teams building document analysis tools, video comprehension pipelines, or agentic interfaces that need to read screens and reason about visual content in a single model.
Quick Info
Powered by- Provider
- Jiekou.AI
- Model key
- zai-org/glm-4.5v
- Release date
- Jan 1, 2026
- Last updated
- Jan 1, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.60
- Output token cost
- $1.80
Limits
- Output tokens
- 16,384 tokens
- Context window
- 65,536 tokens
Latest news about GLM 4.5V
No articles yet. Fetch the latest news to show it here.