Currently listed through these providers:
Model details
GLM 4.5 FP8
GLM-4.5-FP8 is a quantized deployment variant of the GLM-4.5 family, which was built from the ground up as an intelligent agent foundation model. The base architecture is a Mixture-of-Experts design with 355 billion total parameters and 32 billion activated per token, enabling efficient inference without sacrificing capacity. The model's defining feature is a hybrid reasoning approach that can switch between deliberate thinking mode and direct response mode depending on the task, making it adaptable across agentic workflows. The FP8 quantization brings the model into a more manageable footprint for practical deployment while retaining the quality of the full-precision version.
The model was developed through comprehensive post-training that combined expert model iteration with reinforcement learning techniques. It underwent multi-stage training on a dataset of 23 trillion tokens, building the kind of broad capability that agent applications demand. Benchmark results show this training investment pays off: the model achieves 91.0% on AIME 24 reasoning tasks, 70.1% on the TAU-Bench agent evaluation, and 64.2% on the SWE-bench Verified coding benchmark. These scores place it 3rd overall among all evaluated models and 2nd specifically on agentic benchmarks, with much fewer active parameters than several competitors. The compact companion version GLM-4.5-Air, with 106 billion total parameters and 12 billion activated, extends the family for resource-constrained scenarios.
Quick Info
Powered by- Provider
- submodel
- Model key
- zai-org/GLM-4.5-FP8
- Release date
- Jul 28, 2025
- Last updated
- Jul 28, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.20
- Output token cost
- $0.80
Limits
- Output tokens
- 131,072 tokens
- Context window
- 131,072 tokens
Latest news about GLM 4.5 FP8
No articles yet. Fetch the latest news to show it here.