Currently listed through these providers:
Model details
GLM 5.2 Fast
GLM 5.2 Fast is the high-throughput counterpart to Zhipu AI's flagship GLM 5.2, designed for workflows where output speed matters as much as quality. It keeps the standard version's capabilities in logical reasoning, long-form comprehension, and code generation, but applies inference acceleration so that generation runs noticeably faster. In third-party hosting benchmarks, the optimized deployment has reached a peak of around 446 tokens per second on Artificial Analysis and has been consistently ranked among the fastest providers on OpenRouter throughput leaderboards, suggesting real-world latency gains rather than just marketing claims.
The practical appeal of GLM 5.2 Fast is its combination of a very large context window with the throughput needed for interactive products. It is well suited to real-time dialogue, multi-turn agent calls, and streaming code generation, where long context and low time-to-first-token both shape the user experience. Open weights mean teams can self-host or run through optimized inference providers, and the model supports tool calling and structured outputs, making it a flexible engine for production assistants, retrieval-heavy pipelines, and developer tooling that benefits from speed without giving up reasoning depth.
Quick Info
Powered by- Provider
- Baseten
- Model key
- zai-org/GLM-5.2-Fast
- Release date
- Jun 13, 2026
- Last updated
- Jun 13, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $2.10
- Output token cost
- $6.60
Limits
- Output tokens
- 262,144 tokens
- Context window
- 1,048,576 tokens
Transparent token rates
Compare GLM 5.2 Fast pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about GLM 5.2 Fast
No articles yet. Fetch the latest news to show it here.