Currently listed through these providers:
Model details
GLM-5.2
GLM-5.2 is positioned as a flagship open-weight model purpose-built for long-horizon work, particularly the extended trajectories produced by coding agents. According to Z.ai's launch announcement, it represents a substantial step up from its predecessor GLM-5.1 in sustaining quality across long, messy agent sessions rather than simply accepting more tokens, and it does so on a solid 1M-token context window that remains usable in practice. The model ships under a permissive MIT open-source license with no regional access limits, making it attractive for teams that want to self-host or fine-tune without licensing friction. For practitioners, this combination signals a fit for autonomous coding agents, repository-scale refactors, and any workflow where the model must keep coherent state across very long prompts without losing track of earlier decisions.
On the technical side, GLM-5.2 pairs its expanded context with a notable architectural refinement called IndexShare, which reuses a single indexer across every four sparse attention layers and reportedly reduces per-token compute by roughly 2.9× at full 1M length, alongside an improved multi-token prediction layer that lifts speculative-decoding acceptance length by up to 20%. Z.ai also highlights stronger coding capabilities paired with selectable thinking-effort levels, letting developers trade latency against depth of reasoning depending on the task. Together these traits frame GLM-5.2 as a practical choice for latency-sensitive agent pipelines that still demand sustained reasoning quality over long contexts, where the open weights and improved inference efficiency together lower the cost of running frontier-style behavior in production.
Quick Info
Powered by- Provider
- Vultr
- Model key
- zai-org/GLM-5.2-FP8
- Release date
- Jun 13, 2026
- Last updated
- Jun 13, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.85
- Output token cost
- $3.10
Limits
- Output tokens
- 131,072 tokens
- Context window
- 393,216 tokens