GLM-5.2 is Z.ai's flagship release positioned for long-horizon coding and agentic work, succeeding GLM-5.1 as a substantial leap in sustained task capability. The model is distributed under an MIT open-source license with no regional restrictions, with weights hosted under the zai-org organization on GitHub and HuggingFace. Z.ai introduced GLM-5.2 alongside the IndexShare architecture, which reuses a shared indexer across four sparse attention layers to cut per-token compute at long contexts, alongside an improved multi-token prediction layer that extends speculative decoding acceptance length. The combination targets practical engineering use, where long, messy coding-agent trajectories must hold quality across the full session rather than merely accept more tokens.
Beyond raw context length, GLM-5.2 emphasizes advanced coding with flexible thinking effort levels, letting developers trade latency against reasoning depth based on task demands. The model later served as the base for GLM-5.3, with that successor's gains in complex coding, vulnerability discovery on CyberGym, and open-source benchmark performance on Terminal Bench 3.0 and Agents' Last Exam coming entirely from post-training, suggesting a strong underlying foundation in GLM-5.2 suitable for downstream specialization. For teams building autonomous coding pipelines and long-running research or refactoring agents, GLM-5.2 offers an open-weights platform designed to remain reliable across extended sessions while remaining tunable through effort controls and open access.