Zhipu AI
BigGo Finance reports that GLM-5.3-FlashX was released on September 18, 2026, boosting inference speed from the prior 30-50 tokens per second to a peak of 200 tokens per second, a roughly five-to-six-fold improvement. The high-speed tier uses a distinct API identifier and is billed separately at approximately 2.5 times The speed gains are attributed to an inference cluster built on more than 100,000 Chinese-made AI accelerators, with InfraAgent, powered by GLM-5.3, participating in infrastructure optimization in what Zhipu describes as China's first publicly disclosed production recursive self-improvement deployment. The report also