Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Tempr Gateway logo

Model details

GLM-4.6

GLM-4.6 is the flagship text-to-text generation model from Z.ai, positioned as the successor to GLM-4.5 within the broader GLM family. According to Z.ai's official release notes, the model is designed around three core pillars: agentic task execution, advanced reasoning, and real-world coding assistance. Its intended use spans long-context assistant workflows, agent frameworks, and developer-oriented tooling where extended context and tool use during inference are required.

A defining upgrade over GLM-4.5 is the expanded context window, which grew from 128K to 200K tokens, allowing GLM-4.6 to handle more complex agentic tasks that depend on large document or repository context. The model also shows clear improvements in reasoning, supports tool use during inference, and delivers stronger performance in popular coding agents such as Claude Code, Cline, Roo Code, and Kilo Code, including more polished front-end output. Weights are openly published on Hugging Face under the zai-org organization, making GLM-4.6 well suited for self-hosting, research, and custom agent pipelines where open-weight access matters.

Tempr Gatewayzai/glm-4.6glm

Quick Info

Powered by
Provider
Tempr Gateway
Model key
zai/glm-4.6
Release date
Sep 30, 2025
Last updated
Sep 30, 2025
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.60
Output token cost
$2.20

Limits

Output tokens
131,072 tokens
Context window
204,800 tokens

Transparent token rates

Compare GLM-4.6 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-4.6

Impossibl

Official sourceAnnouncement

Z.ai announced GLM-4.6 on 2025-09-30 as the latest version of its flagship model, expanding the context window from 128K to 200K tokens to handle more complex agentic tasks. The model delivers higher scores on code benchmarks and improved real-world performance in Claude Code, Cline, Roo Code, and Kilo Code, including more visually polished front-end generation. Z.ai reports clear gains over GLM-4.5 across eight public benchmarks covering agents, reasoning, and coding. On the extended CC-Bench evaluation, human evaluators ran GLM-4.6 inside isolated Docker containers on multi-turn tasks spanning front-end development, tool building, data analysis, testing, and algorithms, with the model reaching near parity with Claude Sonnet 4 at a 48.6% win rate. Z.ai also notes GLM-4.6 finishes tasks with about 15% fewer tokens than GLM-4.5, and the full evaluation details plus trajectory data are publicly available on Hugging Face for community research and reproduction.

Z.AI

Coverage

Z.AI has officially released GLM-4.6, the latest model in its GLM series, expanding the context window from 128K to 200K tokens to better support complex agentic tasks. The model is live on the Z.ai Chat interface and Z.ai API platform, and has been open-sourced on Hugging Face under the MIT license for broader developer access. GLM-4.6 delivers notable coding and reasoning improvements, integrating with real-world environments including Claude Code, Cline, and Roo Code. Z.ai reports the model matches Claude Sonnet 4 and Sonnet 4.5 across eight benchmarks such as AIME 25, GPQA, and SWE-Bench Verified, while average token consumption dropped over 30 percent compared to its predecessor GLM-4.5.

Z.AI

CoverageAnalysis

Cirra AI published a technical analysis of GLM-4.6 on October 18, 2025, describing it as Zhipu AI's latest flagship mixture-of-experts language model explicitly designed for agentic tasks and tool usage. The model features a 200K-token context window, a reasoning-capable thinking mode, and native support for structured The analysis highlights GLM-4.6's native function-calling reliability, building on GLM-4.5's 90.6% success rate with architecture and fine-tuning that emphasize chain-of-thought planning, argument double-checking, and rejection of unknown tools. GLM-4.6 can autonomously decide when to invoke external tools such as web

Videos about GLM-4.6

More models around GLM-4.6