Currently listed through these providers:
Model details
GLM-5.2
GLM-5.2 is positioned as a flagship model purpose-built for long-horizon work, marking a substantial leap over its predecessor GLM-5.1 in sustaining quality across extended coding-agent trajectories and reasoning chains. Rather than simply accepting more tokens, the model is engineered to maintain coherence over long, messy agentic sessions, delivering a solid the cataloged API limit context that stays stable during complex multi-step tasks. Its intended use centers on advanced software engineering, agentic workflows, complex reasoning, and large-scale data processing, making it well suited for teams running extended coding agents, deep research pipelines, and production pipelines that exceed typical context windows.
On the architecture side, GLM-5.2 introduces IndexShare, a technique that reuses the same indexer across every four sparse attention layers, reportedly reducing per-token FLOPs by roughly 2.9× at the the cataloged API limit context length and making long-context inference more practical. The model also improves its multi-token prediction layer for speculative decoding, increasing acceptance length by up to 20%, which helps balance performance and latency during generation. Coding capability is strengthened through flexible thinking effort levels, letting users trade latency against depth of reasoning. Released under an MIT open-source license with no regional restrictions, GLM-5.2 is available through Z.ai, the Z.ai Coding Plan, and on Hugging Face, giving practitioners broad access to its long-horizon capabilities.
Quick Info
Powered by- Provider
- Volcengine Ark
- Model key
- glm-5-2-260617
- Release date
- Jun 13, 2026
- Last updated
- Jun 13, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $1.18747
- Output token cost
- $4.15615
Limits
- Output tokens
- 131,072 tokens
- Context window
- 1,000,000 tokens