Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Tempr Gateway logo

Model details

GLM-4.5-Flash

The model overview is being prepared.

Tempr Gatewayzai/glm-4.5-flashglm-flash

Quick Info

Powered by
Provider
Tempr Gateway
Model key
zai/glm-4.5-flash
Release date
Jul 28, 2025
Last updated
Jul 28, 2025
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
98,304 tokens
Context window
131,072 tokens

Latest news about GLM-4.5-Flash

Z.AI

Coverage

VentureBeat's July 28, 2025 launch coverage confirms GLM-4.5-Flash as a free variant in the GLM-4.5 family, explicitly optimized for coding and reasoning tasks. Z.ai positioned it as accessible alongside the flagship GLM-4.5 and lighter GLM-4.5-Air, with sibling ultra-fast inference variants GLM-4.5-X and GLM-4.5-AirX also available. The models were released open-source on Hugging Face and ModelScope. The report notes Z.ai supports inference integration via vLLM and SGLang, giving developers multiple deployment paths for GLM-4.5-Flash. The family operates in dual modes—thinking mode for complex reasoning and tool use, and non-thinking mode for instant responses. This third-party corroboration aligns with the official documentation's positioning of Flash as a free, agent-oriented coding and reasoning model.

Z.AI

CoverageBenchmark

A third-party technical analysis on Baidu Cloud examines GLM-4.5-Flash as a free large language model positioned around a zero-cost, high-performance proposition. The article frames the model as breaking traditional free-tier limits through a hybrid reasoning architecture while maintaining a 128K context window equivalent to roughly 200 pages of input. This positions GLM-4.5-Flash as a developer-relevant option for code and data tasks. According to the Baidu technical analysis, GLM-4.5-Flash features automatic mode switching between a Thinking Mode (thinking.type=enabled) for multi-step logical verification in code generation and math, and a Non-Thinking Mode (thinking.type=disabled) that targets under-800ms latency for real-time chat and simple Q&A. The analysis also reports benchmark figures of 27% accuracy improvement on 100,000-line code analysis tasks in Thinking Mode and 120 requests-per-second throughput in basic Q&A for Non-Thinking Mode, though these numbers are promotional and unverified against primary Z.ai documentation.

Z.AI

Official sourceDocumentation

Z.AI's official developer documentation explicitly names GLM-4.5-Flash as a member of the GLM-4.5 model serial, describing it as a free variant offering strong performance for reasoning, coding, and agents. The model shares the GLM-4.5 family's 128K-token context length and supports hybrid reasoning modes, toggled via the thinking.type parameter with dynamic thinking enabled by default. Maximum output reaches 96K tokens. The documentation lists GLM-4.5-Flash alongside siblings GLM-4.5, GLM-4.5-Air, GLM-4.5-X, and GLM-4.5-AirX, with Flash positioned as the accessible free-tier option in the lineup. It supports streaming output, function calling, context caching, and structured output for agent-oriented applications. Input is text-only with text output, reflecting its optimization for reasoning and coding workloads.

Videos about GLM-4.5-Flash

More models around GLM-4.5-Flash