Currently listed through these providers:
Model details
GLM-5.2
GLM-5.2 was positioned as a generation focused on sustained long-horizon work, combining a solid million-token context window with coding capabilities tunable through multiple thinking effort levels that let users trade latency against performance. Architectural innovations included IndexShare, which reuses the same indexer across four sparse attention layers and reduces per-token FLOPs by roughly 2.9× at the full context length, alongside improvements to the multi-token prediction layer that lifted speculative-decoding acceptance length by up to 20%. On standard coding benchmarks, GLM-5.2 was described as the strongest open-weights contender at the time of its launch, with its sparse-attention and MTP refinements serving as the foundation for later gains.
Because GLM-5.3 shares the same base model and inherits its post-training lineage, GLM-5.2 effectively serves as the underlying engine that the newer variant post-trains on top of, with each successive release adding capability rather than replacing the architecture. The model attracted early independent attention through third-party review coverage shortly after release, signaling real-world evaluation of its coding and long-context behavior. For practitioners, this means GLM-5.2 remains a practical fit when an open, flexible-effort reasoning model with a very long context is needed, especially where downstream tuning or further post-training is anticipated.
Quick Info
Powered by- Provider
- Kilo Gateway
- Model key
- z-ai/glm-5.2
- Release date
- Jun 13, 2026
- Last updated
- Jun 13, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $1.40
- Output token cost
- $4.40
Limits
- Output tokens
- 262,144 tokens
- Context window
- 1,048,576 tokens