Currently listed through these providers:
Model details
GLM 4.6V Flash (Free)
GLM-4.6V-Flash is a compact 9-billion parameter vision-language model built for efficient multimodal reasoning. Developed by Zhipu AI (Z.ai), it represents the lightweight end of the GLM-4.6V series, designed specifically for applications where latency and resource constraints matter. The model's defining innovation is its native function calling capability embedded directly into the vision-language architecture—a first for visual models—which allows it to move seamlessly from interpreting images and video to triggering executable actions like search, cropping, or chart analysis. This design makes it particularly suited for frontend automation and scenarios requiring rapid, real-time responses.
The model was trained with an expanded 128K-token context window, enabling it to process substantial multimodal inputs—images, documents, charts, video frames, and text—within a single conversation turn. GLM-4.6V-Flash achieves state-of-the-art visual understanding accuracy within its parameter class, positioning it as a practical choice for developers building multimodal agents that need to perceive visually and act decisively in business workflows.
Quick Info
Powered by- Provider
- ZenMux
- Model key
- z-ai/glm-4.6v-flash-free
- Release date
- Dec 8, 2025
- Last updated
- Dec 8, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 128,000 tokens
- Context window
- 200,000 tokens