Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Tempr Gateway logo

Model details

GLM-5-Turbo

GLM-5-Turbo is designed for agent-driven work that unfolds over long execution chains. The model emphasizes decomposing complex instructions, maintaining stability across extended tasks, and coordinating tools, scheduled actions, and persistent execution rather than handling only short, isolated requests.

Its practical fit is therefore in applications such as autonomous coding agents, multi-step automation, and other workflows that need sustained reasoning, tool use, and reliable continuation over time. The available benchmark evidence highlights strong GPQA performance, while the model’s broad context supports carrying substantial instructions and task history through extended runs.

Tempr Gatewayzai/glm-5-turboglm

Quick Info

Powered by
Provider
Tempr Gateway
Model key
zai/glm-5-turbo
Release date
Mar 16, 2026
Last updated
Mar 16, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$1.20
Output token cost
$4.00

Limits

Output tokens
131,072 tokens
Context window
200,000 tokens

Transparent token rates

Compare GLM-5-Turbo pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-5-Turbo

Impossibl

CoverageRelease Notes

Opper's third-party release tracker lists Z.ai models newest-first and confirms GLM-5-Turbo was released on March 15, 2026, with a 203K context window, $1.20 per million input tokens and $4.00 per million output tokens, and an Artificial Analysis intelligence score of 27. The tracker also documents the broader Z.ai cad The page attributes its release dates and intelligence scores to Artificial Analysis, cross-checked against vendor announcements, and provides an RSS feed for new releases. Pricing shown here differs from VentureBeat's OpenRouter figures, illustrating the gateway-vs-vendor distinction; this aggregator is useful for dat

Z.AI

Coverage

Atlas Cloud announced GLM-5-Turbo availability on its unified API, describing it as developed by Zhipu AI (Z.ai) and tailored for OpenClaw use cases, marking a shift to a closed-source release with higher runtime efficiency than GLM-5 at lower per-call cost. The model supports a context window of up to 200K tokens and delivers improvements in tool use, instruction following, multi-step workflows, and persistent task execution. Atlas Cloud's abstract claims data analysis capabilities comparable to Claude Opus 4.6 and improved performance over GLM-5 in automation and information-processing tasks. The blog notes the model was informally tested under the codename "Pony-Alpha-2" prior to release, framing GLM-5-Turbo as a deployment-ready option for complex business automation, long-document analysis, and software development workflows. Atlas Cloud positions its multi-model routing as enabling cost-effective deployment for GLM-5-Turbo across these enterprise scenarios. The article confirms dynamic reasoning mode support alongside the core agent-execution features.

Z.AI

Coverage

302.AI published a real-world benchmark of GLM-5-Turbo on March 18, 2026, framing it as the "native execution engine for the OpenClaw ecosystem" following Zhipu AI's March 16 release. The model was specialized from early stages for OpenClaw scenarios including environment deployment, development, and analysis, moving beyond conversational commands toward complex long-chain execution. Test results highlighted enhanced tool-calling stability aimed at zero-error execution in complex long-task workflows. The benchmark emphasized four core capabilities: precision tool calling for complex long-task chains, superior instruction decomposition for multi-level long-link tasks, temporal awareness optimized for timed and continuous tasks, and high-frequency processing for high-throughput long-link scenarios. The piece positions GLM-5-Turbo as addressing continuity interruptions common in complex workflows where general-purpose models struggle. 302.AI characterized the release as a shift in AI applications from "chatting" to "executing" tasks autonomously.

Z.AI

Coverage

Z.ai introduced GLM-5-Turbo on March 16, 2026, positioning it as a proprietary variant of its open-source GLM-5 model tuned for agent-driven workflows and OpenClaw-style tasks like tool use, long-chain execution, and persistent automation. The model offers roughly a 202.8K-token context window with 131.1K max output tokens, emphasizing fast inference and stability across extended agent tasks. Z.ai also integrated it into its GLM Coding subscription product, with Pro subscribers receiving access in March and Lite subscribers in April 2026. API pricing for GLM-5-Turbo was listed at $0.96 per million input tokens and $3.20 per million output tokens via third-party provider OpenRouter, making it approximately $0.04 cheaper per total combined cost than its GLM-5 predecessor at one million tokens. The release note describes deep optimization for real-world agent workflows involving long execution chains, including improvements in complex instruction decomposition, tool use, scheduled execution, and task stability. Z.ai is also accepting early-access enterprise applications ahead of broader availability.

Z.AI

Coverage

Zhipu AI released GLM-5-Turbo on March 16, 2026, as a foundation model optimized during training for OpenClaw agent scenarios. The model strengthens tool invocation, instruction following, scheduled and persistent task execution, and long-chain workflows, and ranked first among domestic models on the company's proprietary ZClawBench benchmark. Zhipu AI also introduced OpenClaw Packages subscription plans built on GLM-5-Turbo and released ZClawBench, an end-to-end evaluation benchmark for OpenClaw agent scenarios covering environment setup, coding, information gathering, data analysis, and content creation. The model is available to developers and enterprise users via the BigModel.cn platform.

Z.AI

Official sourceDocumentation

Z.AI's developer documentation positions GLM-5-Turbo as a foundation model deeply optimized for the OpenClaw scenario, with training-stage enhancements targeting tool invocation, command following, timed and persistent tasks, and long-chain execution. The model lists a 200K context length and 128K maximum output tokens across text input and output modalities. The documentation highlights a capability matrix including thinking modes, real-time streaming output, function calling, context caching, structured JSON output, and MCP integration. These features enable tool orchestration, long-conversation performance optimization, and flexible expansion through external MCP tools and data sources for OpenClaw workflows.

Videos about GLM-5-Turbo

More models around GLM-5-Turbo