Currently listed through these providers:
Model details
GLM 4.7
GLM 4.7 is a large-scale open-weights language model designed with a focus on intelligent agents, coding, and multi-step reasoning. Unlike frontier models that chase broad AGI capabilities, GLM 4.7 takes a pragmatic approach—it's engineered as a workhorse for developers and teams that need strong code generation, tool use, and agentic task execution without prohibitive computational costs. The model unifies core coding performance with terminal-based automation and vibe coding for UI and web generation, making it versatile across software development workflows. Its design supports thinking-before-acting behavior, which proves critical for complex agent frameworks like Claude Code, Kilo Code, Cline, and Roo Code, where deliberation improves outcomes on intricate tasks.
The GLM-4.7 family follows a deliberate scaling strategy that balances capability with efficiency. The flagship variant carries 355 billion total parameters with 32 billion active parameters per forward pass, while the Flash variant streamlines to 30 billion total with roughly 3 billion active—a dense alternative to massive proprietary systems. Released by Z.ai, GLM 4.7 represents a meaningful step beyond its GLM-4.6 predecessor, with benchmark gains including a 12.4% improvement on the HLE Humanity's Last Exam reasoning benchmark, a 5.8% jump on SWE-bench coding tests, and significant advances in multilingual coding and tool-use scenarios. The combination of open weights, OpenAI-compatible API access, and predictable pricing positions GLM 4.7 as a practical choice for teams building intelligent agents or automating complex coding workflows without sacrificing frontier-level performance.
Quick Info
Powered by- Provider
- Vercel AI Gateway
- Model key
- zai/glm-4.7
- Release date
- Dec 22, 2025
- Last updated
- Dec 22, 2025
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.60
- Output token cost
- $2.20
Limits
- Output tokens
- 120,000 tokens
- Context window
- 200,000 tokens
Transparent token rates
Compare GLM 4.7 pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.