Currently listed through these providers:
Model details
GLM-4.7
GLM-4.7 is built on a Mixture-of-Experts architecture with 355 billion total parameters, activating 32 billion for each token. This sparse design allows the model to handle specialized tasks without firing all pathways simultaneously, giving it computational efficiency comparable to smaller dense models while maintaining deep, targeted expertise. The architecture supports a 200,000-token context window and a 128,000-token output capacity, enabling it to process entire codebases or generate extensive software frameworks in a single pass. Rather than chasing general-purpose benchmarks, GLM-4.7 was designed as a coding partner tailored for intelligent agents, terminal workflows, and multilingual development scenarios.
The model advances its predecessor through gains across agentic benchmarks, including a 5.8% improvement on SWE-bench for software engineering tasks and a 12.4% leap on the HLE (Humanity's Last Exam) reasoning benchmark. It introduces interleaved thinking capabilities, allowing it to reason before acting when using tools or browsing the web. GLM-4.7 also brings noticeable improvements to frontend generation through its "Vibe Coding" enhancements, producing cleaner, more modern web pages with better layout accuracy. Being fully open-weight, it invites the community to run it locally, fine-tune it for domain-specific workflows, or integrate it into agentic pipelines via OpenAI-compatible APIs. This combination of open accessibility, agent-ready design, and strong coding performance positions GLM-4.7 as a practical foundation for teams building AI-powered development tools.
Quick Info
Powered by- Provider
- Moark
- Model key
- GLM-4.7
- Release date
- Dec 22, 2025
- Last updated
- Dec 22, 2025
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $3.50
- Output token cost
- $14.00
Limits
- Output tokens
- 131,072 tokens
- Context window
- 204,800 tokens
Transparent token rates
Compare GLM-4.7 pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about GLM-4.7
No articles yet. Fetch the latest news to show it here.