Pioneer
The Baidu Cloud blog (published September 7, 2026) explicitly names GLM-5.2 and reports an official June 15, 2026 launch with three evolutionary milestones: continuation of the Mixture-of-Experts sparse architecture at 744B total parameters (40B activated), a context window expansion from 200K to 1M tokens, and an open The same article describes GLM-5.2's MoE implementation as comprising 64 specialized subnetworks with dynamic gating that activates 2–4 experts per token, claims a 60% reduction in theoretical FLOPs and inference latency under 300ms on domestic GPU clusters, and frames the 1M-token context expansion as enabling long-do