Neon
DataCamp reports that Z.ai’s GLM-5.3-Flash uses a 320B-A18B mixture-of-experts design, with 320 billion aggregate parameters and 18 billion active per token, and supports a one-million-token context window. The article describes the exact GLM-5.3-Flash variant as natively multimodal and cost-optimized relative to the l The article gives a Terminal-Bench 2.1 score of 84.3 for GLM-5.3-Flash, compared with 85.0 for Claude Opus 4.8 and 87.4 for GPT-5.6 Terra. It also reports a roughly $0.10 blended price per million tokens, about one-tenth of the cited GLM-5.3 blended price, while noting a 306 GiB FP8 checkpoint that makes lightweight lo