Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Together AI logo

Model details

GLM-5

GLM-5 is positioned as a next-generation foundation model aimed at moving software development from casual "vibe coding" toward full agentic engineering. Built on the agentic, reasoning, and coding strengths of its predecessor, it adopts DeepSeek Sparse Attention (DSA) to meaningfully cut training and inference cost while still holding long context together, an important property when a model has to keep hours of work in mind. The architecture is a 744B-parameter mixture-of-experts design that activates roughly 40B parameters per pass, balancing scale with efficiency so the system can be served at production cost while tackling large codebases and multi-step engineering jobs. Multiple thinking modes let developers trade depth for latency, and the model supports streaming, function calling, context caching, and structured output, making it well suited to integration inside modern agent frameworks and IDE-style coding tools.

On the post-training side, the team rebuilt the reinforcement learning stack around an asynchronous infrastructure that decouples generation from training, dramatically improving throughput, and paired it with new asynchronous agent RL algorithms that let the model learn from long, complex interactions rather than short, isolated prompts. These choices, combined with continued training in agentic, reasoning, and coding (ARC) skills, translate directly into the strengths highlighted by the team and ecosystem partners: 77.8% on SWE-Bench Verified, a leading open-source showing on Vending Bench 2 for long-horizon planning, and a coding experience that the developers describe as approaching Claude Opus 4.5 in real programming scenarios. For practitioners, GLM-5 fits best where a single model has to plan, debug, refactor, and run end-to-end across extended sessions, such as autonomous backend construction, deep debugging with iterative self-correction, and architect-level decomposition of system requirements, while remaining an open-weight option that teams can self-host, inspect, and fine-tune.

Together AIzai-org/GLM-5glmdeprecated

Quick Info

Powered by
Provider
Together AI
Model key
zai-org/GLM-5
Release date
Feb 11, 2026
Last updated
Feb 11, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$1.00
Output token cost
$3.20

Limits

Output tokens
131,072 tokens
Context window
202,752 tokens

Transparent token rates

Compare GLM-5 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-5

Together AI

CoverageBenchmark

The kingy.ai article details the GLM-5.3 open-weight release as of August 28, 2026: the official FP8 repository is 755.7 GB across 141 weight shards, while the BF16 version is approximately 1.5 TB across 282 shards, with Z.ai's vLLM deployment recipe requiring an eight-accelerator server. The piece notes that the previ For practical guidance, the article recommends the API as the rational first step for most teams, with self-hosting justified only by control, data location, customization, or sustained datacenter utilization, and GLM-5.3-Flash preferred when native vision, MIT licensing, or a smaller checkpoint matter more than peak c

Together AI

Coverage

Interconnects AI's August 14, 2026 analysis confirms Z.ai's GLM-5.3 announcement as currently available in the Z.ai coding plan, coming soon to its API and to Hugging Face open weights in roughly two weeks. The piece emphasizes that GLM-5.3 is a pure post-training story on the same GLM-5.2 base model, with Z.ai stating Lambert notes the model surpasses Moonshot AI's Kimi K3 on many benchmarks and matches or exceeds Claude Fable 5 and GPT-5.6-Sol on some, all at approximately 750B parameters (roughly a third of Kimi K3's size). The analysis frames Z.ai's strength as post-training rather than the pretraining-heavy approach attributed t

Together AI

Official sourceBenchmark

Together AI's model page documents GLM-5.2 (served under the zai-org/GLM-5 family on Together), described as Z.ai's flagship open reasoning model for agentic software engineering with a 1M-token usable context window and a 131,072-token output cap. The model is a 744B-parameter Mixture-of-Experts backbone with 40B acti The page highlights three core capabilities: long-context agentic coding (5x larger than GLM-5.1's 200K limit), configurable thinking-effort levels exposed directly at the API, and day-one compatibility with eight coding agents including Claude Code, Cline, Roo Code, Goose, and OpenCode via an OpenAI-compatible and Ant

Together AI

CoverageBenchmark

InferenceX's technical dossier covers GLM-5.2 (Z.ai's flagship long-horizon release, Hugging Face repo created 2026-06-16, MIT-licensed open weights, same API pricing as GLM-5.1) with a solid 1M-token context, multi-level thinking effort, an IndexShare architecture change, and an improved MTP layer for speculative deco The same dossier documents GLM-5.3 (announced August 14, 2026) as a post-training-only release on the GLM-5.2 base, marketed for "Frontier Coding with Emergent Cyber Capabilities," claiming open-source SOTA on Terminal-Bench 3.0 and Agents' Last Exam (CLI). GLM-5.3 is text-only, always reasoning, with 1M context, 128K

Together AI

Official sourceRelease Notes

The Together AI changelog entry for August 28, 2026 announces zai-org/GLM-5.3 as a new serverless model with 1,000,000 context length, FP4 quantization, and support for function calling and structured outputs. Pricing is listed at $1.40 input / $4.40 output / $0.26 cached input per 1M tokens, giving developers a concre The same changelog entry from August 31, 2026 adds Qwen/Qwen3.8-Flash to serverless at $0.15 input / $0.47 output per 1M tokens, while an August 27 update covers a beta Billing usage API (GET /billing/usage), CLI/Python SDK file upload progress, Slurm cluster kubeconfig downloads, legacy API key regeneration, and depre

Videos about GLM-5

More models around GLM-5