Currently listed through these providers:
Model details
GLM 4.7
GLM 4.7 is a large-scale Mixture-of-Experts language model built by Zhipu AI for real-world intelligent agent scenarios. Its sparse MoE architecture activates only a fraction of its 355 billion total parameters per token, allowing it to punch above its weight class while keeping inference manageable. The model is explicitly designed around "thinking before acting," enabling it to deliberate through complex tasks rather than reacting impulsively. Core strengths center on multilingual agentic coding and terminal-based tasks, with particularly strong gains in software engineering benchmarks and frontend "vibe coding" that produces cleaner, more modern web interfaces. Tool-using capabilities have also seen significant improvements, making it well-suited for agent frameworks like Claude Code, Kilo Code, Cline, and Roo Code.
The model builds on the GLM 4.x series lineage as a direct successor to GLM 4.6, and its OpenAI-compatible interface and open weights on Hugging Face make it straightforward to integrate into existing pipelines. Benchmark data shows consistent double-digit percentage gains over its predecessor across SWE-bench, Terminal Bench 2.0, and the HLE reasoning benchmark. For developers seeking frontier-level capability in an open-weight package, GLM 4.7 fits a practical niche: it brings the long-context and output capacity needed for generating entire software frameworks in a single pass, combined with the agentic tooling that production AI assistants require. The availability of a more compact Flash variant further broadens its appeal for teams with varying hardware constraints.
Quick Info
Powered by- Provider
- ZenMux
- Model key
- z-ai/glm-4.7
- Release date
- Dec 22, 2025
- Last updated
- Dec 22, 2025
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.2911
- Output token cost
- $1.1645
Limits
- Output tokens
- 128,000 tokens
- Context window
- 200,000 tokens
Transparent token rates
Compare GLM 4.7 pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about GLM 4.7
No articles yet. Fetch the latest news to show it here.