Currently listed through these providers:
Model details
GLM5.2-Fast
GLM5.2-Fast sits inside Wafer's serverless lineup alongside other optimized open models, where the provider's stated approach is to take existing open models and serve them at meaningfully higher inference speeds than generic providers. The position in the broader GLM lineage suggests an emphasis on retaining the reasoning and conversational strengths associated with that family while redistributing the model for faster, more cost-efficient response in agentic settings. By packaging it behind a pay-as-you-go endpoint that integrates with established coding agents and chat clients, the offering targets developers who want open-model behavior without standing up their own inference stack.
Because the underlying model is drawn from the open-source GLM family, GLM5.2-Fast is well matched to tool-using coding agents, structured data extraction, and multi-turn workflows where reasoning has to stay coherent at long context. The combination of a very large context window, configurable decoding controls, and tool-calling support makes the model practical for retrieval-heavy pipelines, repo-scale code assistance, and other agent loops that mix instructions with large code or document payloads. For teams running agent harnesses such as Claude Code, Codex, Cline, or Roo Code, the appeal is a familiar model personality and capabilities surface delivered with Wafer's latency-focused serving layer.
Quick Info
Powered by- Provider
- Wafer
- Model key
- glm5.2-fast
- Release date
- Jun 13, 2026
- Last updated
- Jun 13, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $3.00
- Output token cost
- $10.25
Limits
- Output tokens
- 131,072 tokens
- Context window
- 1,048,576 tokens
Transparent token rates
Compare GLM5.2-Fast pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about GLM5.2-Fast
No articles yet. Fetch the latest news to show it here.