Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Baseten logo

Model details

GLM 5.2

GLM 5.2 is a 753B-parameter mixture-of-experts reasoning model designed by Z.ai for long-horizon work, with a solid 1M-token context that holds up across extended coding-agent trajectories rather than just accepting more tokens. The architecture introduces IndexShare, a sparse-attention scheme that reuses the same indexer across every four layers and reduces per-token FLOPs by roughly 2.9x at 1M context, paired with an improved Multi-Token Prediction layer that lifts speculative-decoding acceptance length by up to 20%. Weights ship under an MIT license with no regional limits, making the model a rare flagship-scale open release for teams that want to self-host frontier-class reasoning.

GLM 5.2 is positioned for agent-style workflows spanning requirements through multi-platform deployment, with reasoning-effort controls that trade latency against depth and consistent tool use across long sessions. In independent testing it leads on SWE-bench Pro and MCP-Atlas while trailing on Tool-Decathlon and FrontierSWE, and token pricing runs roughly an order of magnitude below top closed peers, a combination that suits project-level software engineering, complex multi-step automation, and budget-sensitive production deployments where 1M-token context actually pays off.

Basetenzai-org/GLM-5.2glm

Quick Info

Powered by
Provider
Baseten
Model key
zai-org/GLM-5.2
Release date
Jun 13, 2026
Last updated
Jun 13, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$1.40
Output token cost
$4.40

Limits

Output tokens
262,144 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare GLM 5.2 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM 5.2

Baseten

Official sourceAnnouncement

Baseten published a technical write-up on July 29, 2026 detailing how it post-trained vision capabilities onto the otherwise text-only GLM 5.2 open model. Rather than training a full vision tower, the team trained only a two-layer MLP projector of roughly 50 million parameters on top of the existing vision encoder from The resulting model reportedly reaches MMMU-Pro parity with Claude 4.5 Haiku at 55%, and the team notes an emergent generalization effect: GLM 5.2 could identify people who were never explicitly labeled in their training data. The post positions the recipe as a lightweight, reproducible path for adding multimodal input

Baseten

Official sourceAnnouncement

Baseten published an engineering deep dive (last updated June 22, 2026) describing how it built the fastest public API for GLM-5.2, reporting over 280 tokens per second as measured by Artificial Analysis on NVIDIA Blackwell hardware. The optimizations span the full inference stack: shared DSA support for the GLM-5.2 ar For developers choosing where to host or call GLM-5.2, this post documents concrete techniques (NVFP4-from-FP8 quantization, Dynamo-based KV-aware routing, PD disaggregation, MTP speculation) and the latency/throughput results Baseten attributes to them. The headline 280+ TPS figure is vendor-reported and externally ve

Baseten

CoverageBenchmark

GLM 5.2 is a 753B parameter MoE model from Z.ai with a 1M-token context window, MIT open-source license, and API pricing that comes in at roughly 6x less than GPT-5.5. It leads on SWE-bench Pro and MCP-Atlas, trails on Tool-Decathlon and FrontierSWE, and introduces IndexShare, an architectural change that cuts per-toke

CoreWeave

CoverageBenchmark

On June 17, 2026, Z.ai published the GLM 5.2 benchmark scorecard that was withheld at the June 13 launch, alongside MIT-licensed open weights for both zai-org/GLM-5.2 and zai-org/GLM-5.2-FP8 on HuggingFace — arriving earlier than the originally promised "the following week" timeline. GLM 5.2 posts 62.1 on SWE-bench Pro The article frames GLM 5.2 as the first credibly open-weight model to lead an Anthropic or OpenAI flagship on a real-world SWE-bench Pro head-to-head, noting that teams using Claude Code, the Claude Agent SDK, Cursor, or the Vercel AI Gateway now have a frontier-tier open-weight backend they can self-host. It also reca

Lilac

CoverageRelease Notes

ThursdAI's June 2026 monthly roundup lists GLM-5.2 as one of 32 AI releases covered that month, attributing it to Z.ai (Zhipu AI) and categorizing it for developers and coding agents. The aggregator reports concrete model-intrinsic specs: a 753-billion-parameter open Mixture-of-Experts architecture and a 1M-token conte As a secondary podcast/newsletter aggregator, ThursdAI corroborates Z.ai authorship and the 753B open-MoE / 1M-context framing, but its single-source GPQA Diamond figure and other benchmark numbers are not independently verified against a Z.ai release post or model card within this candidate set. The roundup neverthele

CoreWeave

CoverageBenchmark

Z.ai (formerly Zhipu AI / THUDM) released GLM-5.2 on June 13, 2026, as a 744-billion-parameter open-weight Mixture-of-Experts model under the MIT license, activating roughly 40B parameters per token across 384 experts and supporting a 1,000,000-token context window with a 131,072-token maximum output. The article docum Technically, GLM-5.2 introduces IndexShare sparse attention, which the source cites as achieving a 2.9x FLOP reduction at the full 1M context length, along with an improved Multi-Token Prediction speculative decoding path. The model uses a higher activation ratio than Kimi K2.7, and the article walks through local depl

Baseten

Coverage

Simon Willison's blog confirms that Z.ai released GLM 5.2 to their coding plan subscribers on June 13, 2026, followed by the full open weights under an MIT license on June 16, 2026. The model is described as a 753B-parameter, 1.51TB mixture-of-experts with 40 active parameters, text-input only, with a 1M-token context The same review notes GLM 5.2 ranks 2nd on the Code Arena WebDev leaderboard behind Claude Fable 5, and that OpenRouter offers it from nine different providers, almost all charging $1.40 per million input and $4.40 per million output — substantially below GPT-5.5 ($5/$30) and Claude Opus 4.5-4.8 ($5/$25). The post also

Baseten

Official sourceRelease Notes

Baseten announced on June 16, 2026 that Z.ai's GLM 5.2 is available through its Model APIs, served via an OpenAI-compatible endpoint that accepts a Baseten API key, with dedicated deployments offered for larger workloads. The changelog positions GLM-5.2 as Z.ai's flagship model for agentic engineering, targeted at long As a first-party availability signal, this post gives developers a clear starting point: route requests to the Baseten Model APIs OpenAI-compatible endpoint with a Baseten API key, or move to a dedicated deployment for sustained throughput. The entry is dated June 16, 2026, the same day as the wider GLM 5.2 release not

Weights & Biases

Coverage

The U.S. National Institute of Standards and Technology's Center for AI Standards and Innovation (CAISI) published an independent assessment of Z.ai's open-weight GLM-5.2 model on July 8, 2026, roughly three weeks after its June 16, 2026 release. CAISI concluded that GLM-5.2 was probably the most capable open-weight AI On safeguards and security, CAISI found mixed results: GLM-5.2's safeguards permitted assistance with agentic cyber exploit development and blocked fewer sensitive biological questions than reference U.S. models, but it appeared potentially more robust than other evaluated PRC open-weight models against agent-hijacking

Baseten

CoverageBenchmark

OpenRouter lists Z.ai's GLM 5.2 as a live large-scale reasoning model with a 1M-token context window, supporting text input/output and reasoning efforts including "high" and "xhigh" (max reasoning). The page shows a release date of June 16, 2026 and indicates the model is suited for long-horizon agent workflows, projec The OpenRouter marketplace page documents two Baseten-hosted GLM 5.2 endpoints, both labeled "Baseten (ZDR) (fp8)" with Zero Data Retention and fp8 quantization. Each is priced at $1.40 per 1M input tokens and $4.40 per 1M output tokens, with cache reads at $0.14. Reported P50 latency is 1.97s at 79 tps (99.90% uptime)

Lilac

CoverageAnalysis

A Hacker News discussion (916 points, ~80 days old) centered on Artificial Analysis's leaderboard post declaring GLM-5.2 the new leading open-weights model. Commenters provide concrete behavioral observations: GLM 5.2 supports reasoning-effort tiers ("high" and "xhigh," with xhigh mapped to max effort), and at xhigh it The thread frames GLM 5.2 as a significant step up for open weights and getting close to frontier, while flagging reasoning efficiency as the next bottleneck relative to GPT 5.5. It also reinforces Z.ai's open-weights positioning by emphasizing the cost gap (GLM 5.2 expected to undercut Opus 4.8 and GPT 5.5 on price ev

Videos about GLM 5.2

More models around GLM 5.2