GLM-5 marks a significant architectural leap from its predecessor, scaling up to 744 billion total parameters with 40 billion active parameters per token, up from the 355 billion total and 32 billion active seen in GLM-4.5. This growth is paired with expanded pre-training data—28.5 trillion tokens compared to 23 trillion—giving the model richer world knowledge to draw from. To keep deployment practical despite the larger scale, GLM-5 integrates DeepSeek Sparse Attention, which preserves long-context reasoning capabilities while reducing operational overhead. The model is purpose-built for complex systems engineering and multi-stage agentic workflows, with design emphasis on the kind of deep, extended problem-solving that powers production-grade coding agents and long-horizon task execution.
The team developed an internal infrastructure called "slime"—an asynchronous reinforcement learning framework—to address the challenge of scaling RL training for large language models. This approach substantially improves training throughput and enables more granular post-training iterations, bridging the gap between a model's base competence and excellence in real-world tasks. The result is a model that achieves best-in-class performance among all open-source models on reasoning, coding, and agentic benchmarks, narrowing the gap with frontier closed models. Released under the MIT license, GLM-5 is designed to be deployed in production environments, with partner platforms offering optimized inference pipelines, one-click deployment, autoscaling, and tools like LoRA adapters and quantization to balance performance and cost for real workloads.