Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Tempr Gateway logo

Model details

GLM-4.6V

GLM-4.6V is an open-source multimodal vision-language model created by Z.ai and released under a permissive MIT license, designed to accept images, video, and text within a single conversation. According to an independent review, the model is offered in two sizes: a full-scale 106 billion parameter build intended for heavier workloads and a 9 billion parameter Flash edition engineered to run on a single high-end GPU. Both variants retain the same multimodal capabilities, giving teams a choice between maximum capacity and a lighter deployment footprint without licensing trade-offs.

In practical terms, GLM-4.6V is positioned as a vision-language system ready for production rather than purely research demos, with native tool calling and a long context window that lets it reason across images, video frames, and lengthy documents together. The Flash variant's small footprint lowers the barrier for self-hosting, while the larger sibling targets organizations that need stronger reasoning across mixed media at scale. A community-modified derivative on Hugging Face points to an official Z.ai repository path for the Flash model, confirming open-weight distribution and making it straightforward for teams to fine-tune or self-host the variant that best fits their latency, cost, and accuracy needs.

Tempr Gatewayzai/glm-4.6vglm

Quick Info

Powered by
Provider
Tempr Gateway
Model key
zai/glm-4.6v
Release date
Dec 8, 2025
Last updated
Dec 8, 2025
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.30
Output token cost
$0.90

Limits

Output tokens
32,768 tokens
Context window
128,000 tokens

Transparent token rates

Compare GLM-4.6V pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-4.6V

Z.AI

CoverageRelease Notes

Z.ai released the GLM-4.6V vision model on 8 December 2025, according to Opper AI's Z.ai release timeline, positioning it alongside GLM 4.7 as the December 2025 pair of GLM drops. The variant is shown with a 131K-token context window and pricing of $0.30 per million input tokens and $0.90 per million output tokens. Opper assigns it an intelligence score of 22, lower than the text-only GLM 4.7 (22) and well below later 2026 releases such as GLM-5.3 (45). GLM-4.6V was the first model in the GLM lineup to carry the V suffix since GLM-4.5V in August 2025, suggesting Z.ai's continued investment in multimodal variants between major text releases. The 131K context is consistent with Z.ai's broader pattern of around 200K windows for flagship models, though GLM-4.6V's window is smaller than the 200K reported for GLM-4.6. No architectural or vision-benchmark detail for GLM-4.6V is provided in the source beyond context length and pricing.

Eden AI

CoverageBenchmark

GLM-4.6V is described as a large multimodal model released on December 8, 2025 by Z.ai, designed for high-fidelity visual understanding and long-context reasoning across images, documents, and mixed media. The model supports up to 128K tokens (listed as 131K in aggregator metadata), processes complex page layouts and c The model page lists input/output modalities as text and image on both ends, with a listed price of $0.30 per million input tokens and $0.90 per million output tokens. Provider hosting data on the page shows NovitaAI and Z.ai as routing targets with separate latency, throughput, and cache-hit metrics, but those gateway

Z.AI

Official sourceRelease Notes

The official Z.AI developer documentation "New Released" page indexes GLM-4.6V as a dated entry in the vendor's release timeline (December 8, 2025), placing it alongside other Z.ai GLM-family releases such as GLM-5.1, GLM-5.2, GLM-5.3, and GLM-5.3-Flash. This first-party registry serves as primary-source confirmation t While the visible excerpt does not quote the GLM-4.6V bullet text itself, the page's structure (per-entry dated headers with capability summaries for each release) establishes it as the canonical Z.AI documentation surface for GLM-4.6V's release notes. The presence of GLM-4.6V in this first-party index, cross-reference

Z.AI

CoverageBenchmark

The SiliconFlow model-info page documents GLM-4.6V as a Z.ai-developed multimodal vision model created on December 8, 2025 and released under the MIT license. It is described as achieving state-of-the-art visual understanding accuracy within its parameter class and, notably, as the first model to natively integrate Fun According to the page, GLM-4.6V's visual context window has been expanded to 128K, enabling long-video-stream processing and high-resolution multi-image analysis. The page outlines concrete developer-facing use cases that depend on this combination of vision and function calling: scientific data analysis over microscop

Videos about GLM-4.6V

More models around GLM-4.6V