GLM-5.3-Flash is Z.ai's flagship open-weight release in the GLM-5 line, introduced as the series' first natively multimodal model. It carries 320B total parameters with only 18B active per token, a sparse-plus-linear attention hybrid that the team says is the first of its kind within the GLM series, designed to keep long-context serving inexpensive without sacrificing precision. Training was rebuilt around a 30T-token multimodal pre-training corpus, and Manifold-Constrained Hyper-Connections (mHC) were used to improve scaling efficiency, giving the model room to deliver stronger reasoning and coding behavior than its direct predecessor, GLM-5.2, while consuming a fraction of the compute and reaching roughly one-tenth the previous price point.
In real-world evaluations, GLM-5.3-Flash approaches Claude Opus 4.8 on coding and agentic benchmarks, a result the developers highlight as the model's defining strength. It also handles attachments, structured output, reasoning, tool calling, and image inputs alongside text generation. The release itself was unconventional: Z.ai tested the system anonymously as "ox-alpha" on OpenCode and OpenRouter before the public reveal, letting community traffic, including praise from Stripe's Patrick Collison, build the model's reputation ahead of launch. For practitioners, GLM-5.3-Flash fits projects that want frontier-class coding and agentic quality on locally deployable open weights, especially when long-context efficiency and multimodal grounding are priorities.