Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
NEAR AI Cloud logo

Model details

GLM-5.1 FP8

GLM-5.1 FP8 is positioned as Z.ai's next-generation flagship model for agentic engineering, designed to keep working productively over very long task horizons rather than plateauing after early progress. The FP8 release is a sparse Mixture-of-Experts checkpoint, large enough that serving it well becomes a distributed systems problem spanning multiple nodes, and the FP8 quantization is intended to make that deployment more tractable on commodity inference hardware. Its training lineage ties back to the GLM-5 family, with a published technical report accompanying the release, and the model is distributed openly so teams can self-host and integrate it into their own pipelines.

In qualitative use, GLM-5.1 is reported to lead GLM-5 by a wide margin on NL2Repo repository generation and Terminal-Bench 2.0 real-world terminal tasks, and to reach state-of-the-art performance on SWE-Bench Pro. Beyond raw benchmark scores, the design focus is on sustained iteration: breaking ambiguous problems down, running experiments, reading results, identifying blockers, and revising strategy across hundreds of rounds and thousands of tool calls. That combination of long-context retention, extended reasoning loops, and structured tool calling makes the FP8 variant a strong fit for self-hosted agentic coding workflows, especially for engineering teams that need control over weights, latency, and infrastructure rather than relying solely on a hosted API.

NEAR AI Cloudzai-org/GLM-5.1-FP8glm

Quick Info

Powered by
Provider
NEAR AI Cloud
Model key
zai-org/GLM-5.1-FP8
Release date
Mar 27, 2026
Last updated
Mar 27, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$1.40
Output token cost
$4.40

Limits

Output tokens
16,384 tokens
Context window
202,752 tokens

Transparent token rates

Compare GLM-5.1 FP8 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about GLM-5.1 FP8

NEAR AI Cloud

Coverage

The vLLM Recipes page for zai-org/GLM-5.1 documents the official deployment configuration for the model served as GLM-5.1 FP8. It explicitly confirms that both BF16 (1,786 GB) and native FP8 (893 GB) checkpoints are published, with the FP8 path requiring vLLM 0.19.0 or later and the DeepGEMM FP8 GEMM kernels. The page For developers, the recipe notes that thinking mode is enabled by default and can be disabled via chat-template kwargs, and it flags that the Rust frontend can improve throughput and latency under high concurrency while Python should be used for unsupported features. It also calls out that nightly vLLM is required to c

Videos about GLM-5.1 FP8

More models around GLM-5.1 FP8