Currently listed through these providers:
Model details
GLM 5.2 Short Fast
GLM 5.2 Short Fast is a ZhipuAI-built text model offered through Neuralwatt's cloud platform, marketed as the quickest and most energy-efficient member of Neuralwatt's 200K-context pool. According to the provider page, it ships with reasoning turned off, which removes chain-of-thought overhead and makes it well suited to latency-sensitive coding assistants, conversational chat, and straightforward tool-calling workflows that don't require step-by-step deliberation. The Neuralwatt portal explicitly tags "Tool Use" as a supported capability, reinforcing its role as a responsive action-oriented model rather than a deep-reasoning engine.
In practical deployment, GLM 5.2 Short Fast is reached through Neuralwatt's OpenAI-compatible /v1/chat/completions endpoint, with simple Bearer-token authentication via a Neuralwatt API key, and the same model ID is routable through third-party frameworks such as Mastra. Beyond the standard tier, it is also listed as available on Neuralwatt's Flex tier, where latency-tolerant requests are billed at a discount, giving operators a way to trade a little responsiveness for lower cost on background or batch-style workloads. The combination of a compact 200K context, open weights, reasoning-off behavior, and explicit tool-use support positions this variant as a workhorse for production pipelines that need fast, predictable text generation rather than extended contemplation.
Quick Info
Powered by- Provider
- Neuralwatt
- Model key
- glm-5.2-short-fast
- Release date
- Jun 17, 2026
- Last updated
- Jun 17, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $1.45
- Output token cost
- $4.50
Limits
- Output tokens
- 32,000 tokens
- Context window
- 199,984 tokens
Latest news about GLM 5.2 Short Fast
No articles yet. Fetch the latest news to show it here.