Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Ofox logo

Model details

GLM-5.2

GLM-5.2 is a Z.ai coding model built around a 753-billion-parameter mixture-of-experts architecture, making it well suited to long-context software work and tool-driven agent tasks. It is open-weight under the MIT license, allowing organizations with sufficient hardware to self-host rather than depend on a hosted service.

The model’s practical appeal is its combination of repository-scale context, tool calling, and always-on reasoning that can be tuned by effort. It is a strong fit for code review, refactoring, batch generation, and other sustained engineering workflows, especially where teams can use its open weights for deployment control and keep an eye on reasoning-token usage.

Ofoxz-ai/glm-5.2glm

Quick Info

Powered by
Provider
Ofox
Model key
z-ai/glm-5.2
Release date
Jun 13, 2026
Last updated
Jun 13, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$1.40
Output token cost
$4.40

Limits

Output tokens
131,072 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare GLM-5.2 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-5.2

Ofox

Official sourceAnnouncement

This Ofox post primarily covers the GLM 5.3 API, but it confirms a key GLM-5.2 data point by anchoring the family's pricing: Z.ai's table now carries a GLM-5.3 row at $1.40 input, $0.26 cached input, and $4.40 output per million tokens — matching GLM 5.2 and GLM 5.1 line for line. The $0.26 cached input rate is 19% of For GLM-5.2-era developers still on older base URLs, the page warns that anything written in the first 24 hours after the GLM-5.x announcement quoted a base URL that did not ship, and that current Z.ai endpoints expose three protocols (OpenAI Chat Completions, OpenAI Responses, Anthropic Messages) with a documented inc

Ofox

Official sourceAnnouncement

On August 14, 2026, Z.ai published a research post on GLM-5.3, explicitly stating that GLM-5.3 uses the same base model as GLM-5.2 and that every reported gain over GLM-5.2 comes from post-training rather than a new base. The post describes scaling the GLM-5.2 stack: IndexShare for efficient long-context processing, SA The same post frames GLM-5.3 as the most capable open-weights coding model Z.ai has released, claiming a 50% improvement over GLM-5.2 on the in-house Z.ai Code Bench and open-source state-of-the-art results on Terminal Bench 3.0 and Agents' Last Exam. It also reports state-of-the-art performance on CyberGym for vulnera

OrcaRouter

Official sourceAnnouncement

Z.ai has officially introduced GLM-5.2, its flagship model purpose-built for long-horizon tasks and a substantial leap over GLM-5.1, delivering for the first time a solid 1M-token context window that the company says remains reliable under real engineering pressure. The model is released under an MIT open-source licens Architecturally, GLM-5.2 introduces IndexShare, which reuses the same indexer across every four sparse-attention layers to cut per-token FLOPs by 2.9× at a 1M context length, alongside an improved MTP layer that increases speculative-decoding acceptance length by up to 20%. The team substantially expanded 1M-context tr

Ofox

Official sourceComparison

Ofox's integration guide shows how to wire GLM-5.2 into the Cline coding agent as a full agent — file reads, diff proposals, command execution — across the model's full 1M-token context by using Cline's OpenAI Compatible provider (not its Anthropic slot, which is reserved for Claude models). The guide flags two practic On cost, the guide reports GLM-5.2 is roughly 1.6–1.9x cheaper than Claude Sonnet 5 on raw per-token rates, with caching narrowing the gap, and recommends GLM-5.2 when Cline sessions are long, file-heavy, and token-volume-driven, while steering readers away when they rely on Claude-native tool use or cache controls. Th

Ofox

Official sourceComparison

Ofox's cost-modeling comparison puts GLM-5.2 at $1.40 input / $4.40 output per million tokens and GPT-5.5 at $5 / $30 on the same ofox.ai OpenAI-compatible endpoint, yielding a 5.56x blended cost ratio at a 2:1 I/O mix ($2.40 vs $13.33 per million tokens) and a 6.82x ratio on pure output tokens. At 100K requests per da The piece recommends GLM-5.2 for cost-sensitive batch coding agents, long-context refactor work, and output-heavy code generation pipelines where input cost and output cost both favor it, and reserves GPT-5.5 for Codex CLI / Terminal-Bench-heavy agentic workflows (citing 82.7% Terminal-Bench 2.1), latency-sensitive int

OpenRouter

CoverageBenchmark

VentureBeat reported on June 16, 2026 that Z.ai released GLM-5.2, a 753-billion-parameter open-weights LLM targeted at long-horizon autonomous coding and engineering. According to the article, GLM-5.2 became available immediately on Hugging Face, the Z.ai API, and more than 20 third-party coding environments, with ente The article corroborates Z.ai's technical claims: a 1-million-token context window, the new IndexShare attention optimization that reuses one indexer across every four sparse attention layers and reduces per-token FLOPs by 2.9× at maximum context, and an upgraded Multi-Token Prediction layer that boosts speculative-dec

Videos about GLM-5.2

More models around GLM-5.2