DigitalOcean
Discover more about what's new at AWS with Minimax M2.5 and GLM 5 models now available on Amazon Bedrock
Model details
GLM-5 is a large language model from Z.ai (zai-org / Zhipu AI) aimed at complex systems engineering and long-horizon agentic tasks, marking a clear shift from earlier chat-oriented use cases toward multi-step software and tooling work. Compared with GLM-4.5, it scales to 744B total parameters with 40B active at inference, and pre-training data grows from 23T to 28.5T tokens, giving the model a much broader base of knowledge to draw on for extended reasoning and code generation.
A defining architectural change in GLM-5 is the integration of DeepSeek Sparse Attention (DSA), which the creators describe as significantly reducing deployment cost while preserving long-context capacity, making the model more practical to serve in production. The accompanying technical report (arXiv 2602.15763), HuggingFace repository (zai-org/GLM-5), and GitHub repository (zai-org/GLM-5) all point to open-weights distribution, so teams can self-host and fine-tune the model for their own agentic and coding pipelines.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
DigitalOcean
Discover more about what's new at AWS with Minimax M2.5 and GLM 5 models now available on Amazon Bedrock
DigitalOcean
Chinese AI company Zhipu AI released GLM-5, a 744-billion-parameter open-source model that rivals Claude Opus 4.5 and GPT-5.2 on coding and agent benchmarks.
DigitalOcean
The GLM-5 represents a shift in AI development from ‘vibe coding’ to ‘agentic engineering’ to generate an enhanced performance.
DigitalOcean
LayerLens ran 20 evaluation batches of base GLM-5 across 13 benchmarks on its Stratix harness over a 24-day window spanning February and March 2026. Reported headline scores include 97.4% on MATH-500 and 96.95% on Human Evaluation, placing GLM-5 among the strongest math reasoners LayerLens had tested. On MATH-500 promp The LayerLens re-evaluation surfaced notable non-monotonic behavior: a 12-point regression on Humanity's Last Exam (from 22.4% down to 10.4%) and a 6.66-point improvement on AIME 2025 (from 86.7% up to 93.3%) within the same 24-day window. Failure-mode analysis pointed to arithmetic errors in word problems, constraint