Venice AI
A detailed third-party breakdown confirms GLM-5.3-Flash launched on August 26, 2026, as a 320B mixture-of-experts activating 18B parameters per token, with MIT-licensed weights on Hugging Face from day one. It is the first natively multimodal model in the GLM-5 series and the first built on a new base since GLM-5.2, wi Pricing is set at $0.15 per million input tokens and $0.50 output, compared with $1.40 and $4.40 for the full GLM-5.3 flagship. Notably, GLM-5.3's own 744B open weights remain missing — Z.ai's Hugging Face organization has no GLM-5.3 repository — making GLM-5.3-Flash the only open-weight model developers can access tod