Currently listed through these providers:
Model details
GLM 5.3 Thinking
GLM 5.3 Thinking is a post-training refinement of the GLM 5.2 base model rather than a fresh pre-training run. Z.AI focused on scaling the post-training stage with additional environments and compute aimed at long-horizon, executable engineering workflows, which is reflected in the way the Thinking variant behaves at inference time: reasoning is always enabled and exposes configurable reasoning effort levels (low, high, max), making the model most useful when chain-of-thought is desired for multi-step tasks. Because the underlying weights and architecture carry over from 5.2, organizations already familiar with GLM 5.2's deployment characteristics should find the transition mostly an exercise in retraining their prompt and routing logic around always-on thinking.
In practical terms, the model targets difficult, agentic coding and security work where extra deliberation pays off. Reported benchmark movement includes large gains on Terminal-Bench 3.0 (jumping to 28.3 from 4.6), DeepSWE v1.1 (66.9 versus 46.2), Agents' Last Exam ALE-CLI (28.5 versus 23.8), and ExploitBench (more than doubling to 54.4), with CyberGym reaching a reported 84.5 state-of-the-art and the Z.AI Code Bench Max improving to 34.5 percent at roughly 75K tokens. Open weights are planned to follow shortly after safety review, which makes the variant attractive for teams that want to self-host a reasoning-tuned model for code generation, tool-using agents, and structured workflows at scale.
Quick Info
Powered by- Provider
- NanoGPT
- Model key
- z-ai/glm-5.3:thinking
- Release date
- Aug 14, 2026
- Last updated
- Aug 14, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $1.00
- Output token cost
- $3.20
Limits
- Input tokens
- 1,048,576 tokens
- Output tokens
- 131,072 tokens
- Context window
- 1,048,576 tokens