Currently listed through these providers:
Model details
GLM-5.3 Highspeed
GLM-5.3 Highspeed sits at the top end of the glm family lineup as a speed-focused coding model intended for agentic and development workflows. The Mastra models table confirms its position alongside glm-5.3 and glm-5.3-flash, all sharing a 1.0M-token context window, which signals Zhipu's continued emphasis on long-context support for complex multi-step coding tasks. The Pi configuration page indicates the model is exposed through an OpenAI-compatible endpoint, allowing developers to integrate it using familiar chat completion patterns while still benefiting from Zhipu-specific reasoning controls.
In practical use, the model supports reasoning effort levels and a zai-specific thinking format, which makes it well suited to tool-using agent loops that benefit from explicit planning phases. Its compatibility profile lists zaiToolStream as enabled, pointing to streamlined tool streaming for code-execution and function-calling scenarios common in coding assistants. The very large context window combined with generous output headroom lets it hold whole repositories or extended interaction histories, making it a natural fit for sustained coding sessions and long-running development agents rather than short conversational exchanges.
Quick Info
Powered by- Provider
- Zhipu AI Coding Plan
- Model key
- glm-5.3-highspeed
- Release date
- Aug 14, 2026
- Last updated
- Aug 14, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 131,072 tokens
- Context window
- 1,000,000 tokens