Z.AI Coding Plan
Z.AI introduced GLM-5.3-Flash on 2026-08-26 as the first natively multimodal model in the GLM-5 series, available through the Z.ai Coding Plan. It is a 320B-total/18B-active MoE that the company says outperforms GLM-5.2 across coding and agentic benchmarks at roughly one-tenth the price, scoring 57 on the Artificial An Architectural changes over GLM-5 include a hybrid sparse + linear attention design to cut long-context serving cost and Manifold-Constrained Hyper-Connections (mHC) for better scaling efficiency, alongside a 30T-token multimodal pre-training corpus. Total parameters sit near GLM-4.5 (320B vs. 355B) but activated parame