Currently listed through these providers:
Model details
GLM-5.2 (EU)
GLM-5.2 is positioned as a flagship model for long-horizon coding and agentic workflows, succeeding the earlier GLM-5.1 lineage with a one-million-token context window designed to keep large repositories and multi-step technical tasks coherent. It exposes several thinking effort levels, letting callers trade deeper reasoning against latency depending on the job, and ships under an MIT license that makes the weights openly available for self-hosted or redistributed deployment. Architectural refinements over the previous generation focus on lowering the cost of inference over long contexts while improving speculative decoding, which together support sustained work on code-heavy projects without paying an outsized price for distant tokens.
In practice the model is aimed at engineers and teams building agent pipelines that need to hold entire codebases, technical specifications, or long conversational histories in view at once, and to call tools reliably while reasoning about them. The open-weight availability, combined with adjustable reasoning depth, makes it a flexible foundation for both hosted assistants and self-managed deployments where control over the stack matters. Its fit is strongest in scenarios such as sustained repository reasoning, complex multi-step tool use, and other long-context technical workloads where the million-token window and configurable inference depth provide a clear advantage.
Quick Info
Powered by- Provider
- Requesty
- Model key
- glm-5.2@eu
- Release date
- Jun 13, 2026
- Last updated
- Jun 13, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $1.20
- Output token cost
- $4.20
Limits
- Output tokens
- 131,072 tokens
- Context window
- 1,048,576 tokens