Umans AI Coding Plan
Fello AI's coverage confirms Z.ai's August 26, 2026 release of GLM 5.3 Flash and frames it as a distinct 320B-A18B MoE built on a new base, not a distilled flagship. It is the first natively multimodal model in the GLM-5 series, advertises a 1M-token context, and uses a hybrid sparse-and-linear attention design. MIT-li The article highlights the pricing gap between Flash and the flagship: $0.15 per million input tokens and $0.50 per million output for Flash, versus $1.40 and $4.40 for GLM 5.3. It also notes that GLM 5.3's promised 744B open-weight release has not materialised, and Z.ai's Hugging Face organisation has no GLM 5.3 repos