Ofox
TechFlow's September 18, 2026 newsletter reports Zhipu's official launch of GLM-5.3-FlashX, confirming the model was previously previewed to global developers under the codename "Ox Alpha." Zhipu stated the 200 tokens/s speed was achieved through increased infrastructure investment and inference optimization on a computing cluster backed by 100,000 domestically produced chips, with the FlashX API going live under the identifier "GLM-5.3-FlashX." The launch framing positions FlashX as combining intelligence, price, and speed simultaneously for the first time within the GLM-5 series. The article notes rapid prior adoption of the underlying GLM-5.3-Flash among global developers drove Zhipu to scale inference capacity, making FlashX a direct response to enterprise demand for higher-throughput, lower-latency access to the same model weights.