Currently listed through these providers:
Model details
GLM 4.7 Flash Heretic
GLM 4.7 Flash Heretic is a community-created abliterated variant derived from Zhipu AI's GLM-4.7-Flash model, with the "Heretic" designation indicating that refusal mechanisms were removed through aggressive abliteration rather than retraining. The model uses a Mixture-of-Experts architecture totaling 30B parameters, with approximately 3B active during inference, allowing it to run efficiently on consumer hardware when quantized. This open-weights design makes the model accessible for local deployment and customization by developers seeking fewer content restrictions in their workflows.
Practical fit centers on lightweight, high-throughput use cases where uncensored reasoning is acceptable. The architecture enables Q4_K_M quantization that fits within 16GB VRAM, demonstrated on setups like an RTX 5070 Ti using Ollama, making it suitable for real-time conversational applications. Strong multilingual performance, particularly in Chinese, extends its usefulness for diverse language tasks, while built-in capabilities for reasoning, tool calling, and structured output make it well-suited for agentic pipelines and complex prompt workflows that demand transparent, step-by-step thinking.
Quick Info
Powered by- Provider
- Venice AI
- Model key
- olafangensan-glm-4.7-flash-heretic
- Release date
- Feb 4, 2026
- Last updated
- Jun 11, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.07
- Output token cost
- $0.40
Limits
- Output tokens
- 24,000 tokens
- Context window
- 200,000 tokens
Latest news about GLM 4.7 Flash Heretic
No articles yet. Fetch the latest news to show it here.