Currently listed through these providers:
Model details
GLM-5.2
GLM-5.2 is presented by its publisher as a flagship model built specifically for long-horizon tasks, framed as a substantial leap over its predecessor GLM-5.1. The headline positioning centers on sustained quality across extended, messy coding-agent trajectories rather than simply accepting more tokens, and the model is marketed with a solid one-million-token context window that the team emphasizes remains stable under real agent workloads. Alongside the long-context push, GLM-5.2 introduces an advanced coding mode with multiple thinking-effort levels, letting users trade latency against reasoning depth depending on the task.
On the architecture side, Z.ai highlights an "IndexShare" design that reuses the same indexer across every four sparse attention layers, reportedly reducing per-token compute at long context, and an improved multi-token prediction layer aimed at speeding up speculative decoding. The model is released under an MIT license with weights published on Hugging Face and the codebase on GitHub, removing regional access limits and making it well suited for teams that need to self-host, fine-tune, or deeply inspect a coding-oriented agent model with a generous context budget.
Quick Info
Powered by- Provider
- Scaleway
- Model key
- glm-5.2
- Release date
- Jun 13, 2026
- Last updated
- Jun 13, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $1.80
- Output token cost
- $5.50
Limits
- Output tokens
- 16,384 tokens
- Context window
- 256,000 tokens