Bothub
Z.ai released GLM-5.3-Flash on August 26, 2026, shipping it as a 320-billion-parameter mixture-of-experts model that activates just 18 billion parameters per token, with MIT-licensed weights posted to Hugging Face on day one. According to Z.ai's announcement quoted in the piece, it is the first natively multimodal mode API pricing for GLM-5.3-Flash is set at $0.15 per million input tokens and $0.50 per million output tokens, roughly one-ninth the cost of the GLM-5.3 flagship's $1.40/$4.40 rates. The piece also notes that Z.ai's promised open-weight release of the GLM-5.3 flagship's 744B parameters, expected around launch, has not mat