GLM-5.2 is Z.ai's latest flagship model, built from the ground up around long-horizon work, and ships under a permissive MIT license with no regional limits. The defining feature is a solid one-the cataloged API limit that the team trained explicitly for coding-agent scenarios, covering large-scale implementation, automated research, performance optimization, and complex debugging rather than merely accepting more tokens. To make that scale affordable, the architecture adds an indexer reused across every four sparse attention layers, which the creators report cuts per-token compute by roughly 2.9x at full context, alongside an improved multi-token prediction layer for speculative decoding that lifts acceptance length by up to twenty percent.
The release targets agentic coding and tool use as its practical sweet spot. Z.ai reports that GLM-5.2 is the strongest open-source model on standard coding suites such as Terminal-Bench 2.1 and SWE-bench Pro, and on FrontierSWE it trails the leading closed model by only a small margin while ranking as the highest-scoring open entry. Multiple thinking effort levels let callers trade latency for capability, and an anti-hack module plus a critic-based training formulation address reward-gaming and trajectory-compaction challenges that show up when agents run for hours. CAISI's independent assessment found GLM-5.2 to be the most capable open-weight model at release, with overall capability comparable to GPT-5.2, making it a credible backbone for self-hosted coding agents and research assistants that need sustained, multi-hour execution.