Currently listed through these providers:
Model details
glm-4.6
GLM-4.6 is a 355-billion-parameter Mixture-of-Experts model built by Zhipu AI to push open-weight models into frontier-scale territory. What sets it apart is the permissive MIT license—rare at this scale—which lets enterprises deploy, fine-tune, and own the model without vendor lock-in. Designed as a bilingual champion, it ranks as the top domestic model for Chinese and English codebases, making it especially strong for developers working across both languages. Its agentic design emphasizes tool integration and long-context reasoning, enabling it to handle complex multi-step tasks that require keeping track of extended上下文.
The model advances a lineage from earlier GLM generations, achieving clear gains in coding benchmarks, reasoning tasks, and real-world agent performance over its predecessor. Evaluated across eight public benchmarks covering agents, reasoning, and coding, GLM-4.6 reaches near-parity with Claude Sonnet 4 on certain tasks while still lagging behind frontier models in some areas. Its 15% token-efficiency improvement makes self-hosting more cost-effective, and support for FlashAttention-4 accelerates inference on modern GPU hardware. Available through both API and Hugging Face, GLM-4.6 serves developers, researchers, and enterprises seeking open, customizable AI that remains competitive with proprietary alternatives.
Quick Info
Powered by- Provider
- 302.AI
- Model key
- glm-4.6
- Release date
- Sep 30, 2025
- Last updated
- Sep 30, 2025
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.286
- Output token cost
- $1.142
Limits
- Output tokens
- 131,072 tokens
- Context window
- 204,800 tokens
Transparent token rates
Compare glm-4.6 pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about glm-4.6
No articles yet. Fetch the latest news to show it here.