Currently listed through these providers:
Model details
GLM-5.2
GLM-5.2 is positioned as Z.ai's flagship model for long-horizon tasks, representing a substantial leap over its predecessor GLM-5.1 in sustaining extended, multi-step work. Its design centers on a solid 1M-token context that remains usable across long, messy coding-agent trajectories rather than merely accepting more tokens, with a maximum output of 128K tokens and text-only input and output. The model is intended to carry a single task through the full development workflow, from requirements to deployable products across multiple platforms, making project-scale engineering context practical rather than theoretical.
Architecturally, GLM-5.2 introduces the IndexShare sparse attention mechanism, which reuses the same indexer across every four sparse attention layers and reduces per-token FLOPs by a reported 2.9× at 1M context length. Its MTP layer was also improved for speculative decoding, increasing the acceptance length by up to 20%. Coding performance can be tuned through multiple thinking effort levels to balance latency against capability, and the model ships under an MIT open-source license with no regional access restrictions. Together, these features suit teams that need an open-weight long-context model for autonomous coding agents, large codebase refactors, and other engineering tasks that span very long task horizons.
Quick Info
Powered by- Provider
- Abacus
- Model key
- zai-org/GLM-5.2
- Release date
- Jun 13, 2026
- Last updated
- Jun 13, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $1.40
- Output token cost
- $4.40
Limits
- Output tokens
- 131,072 tokens
- Context window
- 1,048,576 tokens
Transparent token rates
Compare GLM-5.2 pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.