Sulat.com
AI models
submodel logo

Model details

GLM 4.5 FP8

GLM-4.5-FP8 is a quantized deployment variant of the GLM-4.5 family, which was built from the ground up as an intelligent agent foundation model. The base architecture is a Mixture-of-Experts design with 355 billion total parameters and 32 billion activated per token, enabling efficient inference without sacrificing capacity. The model's defining feature is a hybrid reasoning approach that can switch between deliberate thinking mode and direct response mode depending on the task, making it adaptable across agentic workflows. The FP8 quantization brings the model into a more manageable footprint for practical deployment while retaining the quality of the full-precision version.

The model was developed through comprehensive post-training that combined expert model iteration with reinforcement learning techniques. It underwent multi-stage training on a dataset of 23 trillion tokens, building the kind of broad capability that agent applications demand. Benchmark results show this training investment pays off: the model achieves 91.0% on AIME 24 reasoning tasks, 70.1% on the TAU-Bench agent evaluation, and 64.2% on the SWE-bench Verified coding benchmark. These scores place it 3rd overall among all evaluated models and 2nd specifically on agentic benchmarks, with much fewer active parameters than several competitors. The compact companion version GLM-4.5-Air, with 106 billion total parameters and 12 billion activated, extends the family for resource-constrained scenarios.

submodelzai-org/GLM-4.5-FP8glm

Quick Info

Powered by
Provider
submodel
Model key
zai-org/GLM-4.5-FP8
Release date
Jul 28, 2025
Last updated
Jul 28, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.20
Output token cost
$0.80

Limits

Output tokens
131,072 tokens
Context window
131,072 tokens

Latest news about GLM 4.5 FP8

No articles yet. Fetch the latest news to show it here.

Videos about GLM 4.5 FP8

More models around GLM 4.5 FP8