LLM Gateway
Z.ai officially released GLM-5.3-Flash on August 26, 2026, as a newly trained model rather than a post-train of GLM-5.2, trained on a 30T-token multimodal corpus. The architecture combines sparse and linear attention with Manifold-Constrained Hyper-Connections, reducing active parameters from 32B to 18B and cutting attention compute roughly 3.0× versus GLM-5.3 with a 4.4× smaller KV-cache. The model serves efficiently at 1M-token context lengths and was anonymously evaluated as Ox Alpha on Chinese AI accelerators before launch. Z.ai reports GLM-5.3-Flash outperforms GLM-5.2 across coding and agentic benchmarks, scoring 84.3 on Terminal Bench 2.1, 63.4 on DeepSWE v1.1, 56.3 on NL2Repo, 78.4 on Toolathlon Verified, 48.8 on AutomationBench v1.0.6, 26.3 on Agents' Last Exam, 55.3 on HLE with Tools, and 1773 on GDPval-AA v2. The IndexPool mechanism further compresses cached keys at million-token contexts, supporting long-horizon agent workloads. Weights ship under MIT, with vLLM and SGLang recipes available for self-hosting.