Currently listed through these providers:
Model details
GLM 5.2 Fast
GLM 5.2 Fast sits inside the GLM family from Z.AI and is positioned as a low-latency text generation variant, surfaced on the Infron marketplace as a newly featured entry under the z-ai/glm-5.2-fast identifier. The listing advertises a text-only interface with streaming output, tool calling, and structured JSON responses, pointing to a Quick Start for text workflows. Infron exposes the model through two service tiers, Flex and Standard, that share the same underlying weights and output but differ in the provider pool they route to, with the Standard tier showing a 72-hour uptime of 100.00% after routing and fallbacks at the time of capture. Per-tier average latency, time-to-first-token, and throughput metrics are tracked by Infron, though the snapshot supplied here reports those values as unmeasured rather than as a settled benchmark.
Practically, the model is aimed at production text workloads that benefit from the GLM lineage, including agentic and tool-using applications where structured JSON output and reliable streaming matter. Z.AI's fast-tier branding suggests a focus on speed-sensitive interactive use cases, while the marketplace's tiered routing gives teams a way to trade variability tolerance for stability depending on whether they pick Flex for batch and dev work or Standard for steady production traffic. Pricing on Infron differs between tiers, with the Standard list price coming in above the Flex list price for both input and cache reads, giving operators a clear cost lever alongside the routing choice. Because the only independent evidence is the marketplace listing, architectural specifics and benchmark results beyond the marketplace's own reliability snapshot are not established here.
Quick Info
Powered by- Provider
- Neuralwatt
- Model key
- glm-5.2-fast
- Release date
- Jun 17, 2026
- Last updated
- Jun 17, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $1.45
- Output token cost
- $4.50
Limits
- Output tokens
- 1,048,560 tokens
- Context window
- 1,048,560 tokens
Transparent token rates
Compare GLM 5.2 Fast pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about GLM 5.2 Fast
No articles yet. Fetch the latest news to show it here.