Z.AI
The Hacker News thread on GLM-4.7-Flash (378 points, 135 comments, submitted by scrlk) captures community testing of the model in OpenCode running local 30B-A3B quantizations via llama.cpp on a 32 GB GPU with 128k context. One commenter reports that Qwen3-Coder had previously given the best results in their workflow of Early user issues documented in the thread include broken tool calling in OpenCode, rapid self-repetition (mitigated by raising llama.cpp's --dry-multiplier to 1.1 or higher), and spelling errors such as class or file-name characters being replaced with "1" or "AGENTS.md" being misread as "AGANTS.md". A later reply ann