Z.AI
Z.AI officially introduced GLM-5.3-Flash on August 26, 2026, describing it as the first natively multimodal model in the GLM-5 series. It is a Mixture-of-Experts model with 320B total parameters and 18B active parameters, built on a newly trained base rather than a post-train of GLM-5.2, and trained on a 30T-token mult The model introduces a hybrid sparse and linear attention architecture combined with Manifold-Constrained Hyper-Connections, which Z.AI credits with sharply reducing long-context serving cost while preserving long-context capability. Compared with the GLM-4.5 series, it nearly halves both activated parameters (18B vs 3