Currently listed through these providers:
Model details
GLM-4.7-Flash
GLM-4.7-Flash is positioned by its distribution on Ollama as the strongest model in the 30B class, framing it as a lightweight option that aims to balance performance and efficiency for practical deployments. The Ollama listing reports roughly 1.3 million pulls, signaling meaningful community uptake for a model pitched at smaller-scale, on-device use cases rather than frontier-scale serving. Its inclusion alongside developer-oriented applications such as Claude Code, Codex App, OpenClaw, Hermes Agent, Codex, and OpenCode suggests the model is being adopted as a local backend for agentic and code-assistant workflows where a compact footprint matters more than absolute capability.
Independent technical activity around GLM-4.7-Flash is already visible: an NVIDIA developer forums thread from January 2026 explicitly requests AWQ and NVFP4 quantized instructions for the model under the DGX Spark / GB10 projects category, reflecting real-world interest in running the weights on accelerated and edge hardware. That request thread, combined with the Ollama integration examples, indicates the model is being treated as a deployable open-weight artifact that practitioners want to compress, fine-tune, and wire into local coding and agent stacks. For teams choosing a text-to-text model in the tens-of-billions parameter range, GLM-4.7-Flash fits scenarios that prioritize fast local inference and tooling integration over maximum reasoning depth.
Quick Info
Powered by- Provider
- Jiekou.AI
- Model key
- zai-org/glm-4.7-flash
- Release date
- Jan 1, 2026
- Last updated
- Jan 1, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.07
- Output token cost
- $0.40
Limits
- Output tokens
- 128,000 tokens
- Context window
- 200,000 tokens
Latest news about GLM-4.7-Flash
No articles yet. Fetch the latest news to show it here.