Currently listed through these providers:
Model details
GLM-4.7 Flash (EmberCloud)
GLM-4.7 Flash is positioned within its ecosystem as the strongest offering in the 30B size class, designed for lightweight local deployment that balances performance with efficiency. Public distribution through the Ollama model library makes it accessible to developers who want to run inference on their own hardware, with documented usage patterns that range from simple curl calls against a local chat endpoint to Python and JavaScript client integrations. The Ollama listing also packages the model for several third-party coding and agent workflows, including Claude Code, OpenCode, and Hermes Agent, suggesting a deliberate focus on developer tooling rather than general conversational use.
The model's local-first orientation is reinforced by its strong traction in community channels, with the Ollama entry reporting roughly 1.5 million downloads and an active update cadence that points to ongoing maintenance of the hosted artifacts. Quantitative tag listings on the Ollama page indicate a text-in, text-out chat profile with a sizeable context window suitable for code repositories and longer technical prompts, while quantized variants are made available to help users fit the model on more modest hardware. Together, these signals describe a practical fit for developers who want a capable, locally runnable assistant for coding, scripting, and agent-style tasks without depending on a hosted API.
Quick Info
Powered by- Provider
- LLM Gateway
- Model key
- embercloud/glm-4.7-flash
- Release date
- Jan 19, 2026
- Last updated
- Jan 19, 2026
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.06
- Output token cost
- $0.40
Limits
- Output tokens
- 131,000 tokens
- Context window
- 200,000 tokens
Latest news about GLM-4.7 Flash (EmberCloud)
Videos about GLM-4.7 Flash (EmberCloud)
More models around GLM-4.7 Flash (EmberCloud)
This exact model name is also listed by 17 other providers.