Currently listed through these providers:
Model details
GLM 5.2 Short
GLM 5.2 Short sits within Neuralwatt's lineup as a "Short" variant of the GLM 5.2 family, distinguished by a deliberately smaller working envelope compared with the flagship 1M-context versions. The catalog positions it for concise generative tasks rather than long-document workloads, and it inherits family-level support for reasoning and tool calling while exposing temperature control for output calibration. Its place in Neuralwatt's tier grid is clear: the same shape and capability profile appear across multiple SKUs (standard, fast, flex), letting users trade latency or cost against the short-context budget without leaving the GLM 5.2 family.
In practice, GLM 5.2 Short is best suited to interactive assistants, short-form drafting, code snippets, and retrieval-augmented workflows whose payloads stay well under the supported envelope, where the reasoning and tool-calling features are most likely to be exercised. The open weights flag makes it attractive for teams that want local or self-hosted experimentation alongside the hosted OpenAI-compatible endpoint at api.neuralwatt.com/v1, while the family variants offer a pragmatic migration path: shifting to GLM 5.2 Short Flex or its fast sibling when budget pressure rises, or stepping up to the full-context GLM 5.2 when the workload grows beyond the short window. For builders already using Mastra or AI SDK clients, the model slots into existing OpenAI-compatible stacks with only an API key change, keeping integration overhead low.
Quick Info
Powered by- Provider
- Neuralwatt
- Model key
- glm-5.2-short
- Release date
- Jun 17, 2026
- Last updated
- Jun 17, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $1.45
- Output token cost
- $4.50
Limits
- Output tokens
- 32,000 tokens
- Context window
- 199,984 tokens
Latest news about GLM 5.2 Short
No articles yet. Fetch the latest news to show it here.