Tencent Coding Plan (China)
Discover more about what's new at AWS with Minimax M2.5 and GLM 5 models now available on Amazon Bedrock
Model details
GLM-5 is positioned as a step beyond conversational coding assistance toward full agentic engineering, designed to handle complex systems work and long-horizon tasks that require sustained planning and tool use. The model scales up substantially from its predecessor, moving to 744B total parameters with 40B active per inference and growing pre-training data to 28.5 trillion tokens, up from 355B parameters with 32B active and 23 trillion tokens in the prior generation. This combination of a sparse Mixture-of-Experts design and a much larger training corpus is what allows GLM-5 to operate as a foundation for multi-step software and reasoning workflows rather than single-turn code completion.
A defining architectural addition is the integration of DeepSeek Sparse Attention, which trims deployment overhead while keeping the model capable of handling long contexts, making the larger parameter count more economical to serve. The team also introduced slime, an asynchronous reinforcement learning infrastructure aimed at faster and more fine-grained post-training iteration, and reporting indicates the model was trained on a cluster of Huawei Ascend 910B accelerators rather than NVIDIA hardware. GLM-5 has been made broadly accessible, including availability through Amazon Bedrock, giving teams a frontier-scale open agentic model option for engineering pipelines that need sustained reasoning, tool use, and temperature-controlled generation.
A provider subscription or plan supersedes token-based pricing for this model.
Tencent Coding Plan (China)
Discover more about what's new at AWS with Minimax M2.5 and GLM 5 models now available on Amazon Bedrock