Currently listed through these providers:
Model details
GLM 5.2 Flex
GLM 5.2 Flex is part of the broader GLM 5.2 family that Neuralwatt hosts as provider-specific variants through its OpenAI-compatible chat completions endpoint. Within that family it sits as the cost-optimized "Flex" tier, sitting below the standard and "Fast" GLM 5.2 SKUs while preserving the same million-token window. It is offered alongside sibling "Short" Flex cuts that trim context to roughly 200K tokens, giving teams a way to trade depth for spend on long-context workloads.
Practically, GLM 5.2 Flex is shaped for agentic and analytical pipelines that need very long inputs without paying top-tier rates: Neuralwatt's catalog lists it with reasoning, tool calling, and temperature control enabled, while structured-output support is not indicated for the variant. The large context and output ceiling make it a fit for retrieval-heavy assistants, codebase-aware workflows, and multi-document synthesis where preserving full material matters more than selecting the premium-tier GLM 5.2 endpoint.
Quick Info
Powered by- Provider
- Neuralwatt
- Model key
- glm-5.2-flex
- Release date
- Jun 17, 2026
- Last updated
- Jun 17, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.9425
- Output token cost
- $2.925
Limits
- Output tokens
- 1,048,560 tokens
- Context window
- 1,048,560 tokens
Latest news about GLM 5.2 Flex
No articles yet. Fetch the latest news to show it here.