Merge Gateway
InferenceX provides a technical deep-dive covering GLM-5 and GLM-5.1, noting that GLM-5 scales from 355B parameters (32B active) to 744B parameters (40B active) with 28.5T pre-training tokens, and that GLM-5.1 is a point release on the same architecture with stronger coding and SOTA SWE-Bench Pro. The article cites the GLM-5.1's distinguishing claim per the coverage is long-horizon durability: sustaining optimization over hundreds of rounds and thousands of tool calls where earlier models exhaust their repertoire early. The article notes GLM-5 integrates DeepSeek Sparse Attention (DSA) to reduce deployment cost while preserving long-