GLM-5.3 Flash, released by Z.ai on August 26, 2026, is the first natively multimodal member of the GLM-5 family and was previewed anonymously on OpenRouter as "Ox Alpha" before its public identification. It is built on a newly trained 320-billion-parameter Mixture-of-Experts base that activates roughly 18 billion parameters per token, combining sparse and linear attention layers with Manifold-Constrained Hyper-Connections to keep long-context serving economical. Training drew on a 30-trillion-token multimodal corpus and the weights ship under an MIT license on Hugging Face, with a public vLLM recipe available for local deployment. A native vision encoder handles images and video through the same token stream as text, so the model can ground code, tool calls, and reasoning in visual evidence without a separate vision pipeline.
GLM-5.3 Flash is positioned as an efficient coding and agent workhorse rather than a trimmed flagship, and independent testing confirms the framing. Artificial Analysis places it at the high end of speed among similarly sized open-weight models while noting that it can be verbose on long reasoning tasks, and Z.ai's published results include 84.3 on Terminal-Bench 2.1, 63.4 on DeepSWE v1.1, and 48.8 on AutomationBench, with an independent GDPval-AA v2 Elo that leads several closed alternatives. The combination of a million-token context, tool calling, and native image and video understanding makes it well suited to repository-scale coding agents, document and spreadsheet workflows, and multimodal inspection pipelines where routing routine steps to a lower-cost model preserves budget for the hardest reasoning calls.