Currently listed through these providers:
Model details
GLM 5.2 Short Fast Flex
GLM 5.2 Short Fast Flex belongs to the GLM family of text models, a lineage that the Neuralwatt catalog splits into parallel Short and long-context branches with Fast and Flex throughput variants. Within that structure, this variant combines the short context profile with the Flex pricing tier and the Fast path, positioning it as a latency-oriented, cost-efficient option for short, high-volume interactions rather than for long-document reasoning. Its placement in the same row as the base GLM 5.2 in the Neuralwatt table suggests it inherits the same underlying architecture while trading throughput and reasoning features for a lower per-token price.
Because the Fast variant disables reasoning mode in the provider's compatibility layer, the model is best suited to agent loops, tool-calling chains, and templated generation where deterministic, low-latency responses matter more than extended internal deliberation. The 199,984-token context and output window is generous for short-form workflows such as command interpretation, structured extraction, and iterative code assistance, while the open-weights availability lets teams self-host or fine-tune the model for proprietary pipelines. Taken together, GLM 5.2 Short Fast Flex is a practical pick for builders who want GLM 5.2 family behavior on a Fast, low-cost budget rather than maximum depth.
Quick Info
Powered by- Provider
- Neuralwatt
- Model key
- glm-5.2-short-fast-flex
- Release date
- Jun 17, 2026
- Last updated
- Jun 17, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.9425
- Output token cost
- $2.925
Limits
- Output tokens
- 32,000 tokens
- Context window
- 199,984 tokens
Latest news about GLM 5.2 Short Fast Flex
No articles yet. Fetch the latest news to show it here.