GLM 5 is a large-scale open-weights language model built around complex systems engineering and long-horizon agentic work. It continues the GLM family lineage, scaling up from the earlier GLM-4.5 design of 355B parameters with 32B active on 23T pre-training tokens to a substantially larger 744B-parameter architecture with 40B active and 28.5T tokens of pre-training data, reflecting a continued bet on raw scale as a path to stronger intelligence efficiency. The model is distributed openly through a public GitHub repository, a Hugging Face release, and a technical report on arXiv, positioning it as a transparent option for teams that want to run or inspect the weights themselves.
A defining architectural choice in GLM 5 is the integration of DeepSeek Sparse Attention, which is designed to lower deployment cost while preserving the ability to handle long contexts, a useful combination for agentic workloads that accumulate large state. To support the post-training side, the team built slime, an asynchronous reinforcement learning infrastructure aimed at improving RL training throughput and enabling more fine-grained iteration, addressing a known bottleneck when scaling RL to frontier-sized models. Together these design decisions support the model's positioning as a strong open-source performer for reasoning, coding, and agentic tasks, and make it a practical fit for engineering teams building end-to-end systems rather than one-off completions.