Z.AI
VentureBeat's July 28, 2025 launch coverage confirms GLM-4.5-Flash as a free variant in the GLM-4.5 family, explicitly optimized for coding and reasoning tasks. Z.ai positioned it as accessible alongside the flagship GLM-4.5 and lighter GLM-4.5-Air, with sibling ultra-fast inference variants GLM-4.5-X and GLM-4.5-AirX also available. The models were released open-source on Hugging Face and ModelScope. The report notes Z.ai supports inference integration via vLLM and SGLang, giving developers multiple deployment paths for GLM-4.5-Flash. The family operates in dual modes—thinking mode for complex reasoning and tool use, and non-thinking mode for instant responses. This third-party corroboration aligns with the official documentation's positioning of Flash as a free, agent-oriented coding and reasoning model.