NaN
GLM-5.3-Flash is a sparse mixture-of-experts model from Z.ai with 320B total parameters and 18B active per token, released on August 26, 2026 in FP8 and BF16 checkpoints under an MIT license, with weights published on Hugging Face and API pricing set at $0.15 per million input tokens. According to the LumaDock explaine On Z.ai's launch table, GLM-5.3-Flash posts 63.4 Pass@1 on DeepSWE v1.1 (versus 46.2 for GLM-5.2), 84.3 on Terminal-Bench 2.1 (Claude Opus 4.8 at 85.0 and GPT-5.6 Terra at 87.4), and 55.3 on Humanity's Last Exam with tools, and Z.ai claims a hybrid linear-plus-sparse attention design that delivers roughly 3x less atten