Cortecs
The 30B model achieves 59.2% on SWE-bench Verified while running at 80+ tokens per second on MacBooks.
Model details
GLM-4.7-Flash is designed as a pragmatic, high-performance alternative to larger, more complex systems. While its flagship counterpart utilizes a massive Mixture-of-Experts architecture, this model employs a streamlined, dense 30-billion parameter structure. This design choice ensures that every parameter is utilized for every token processed, resulting in highly predictable inference behavior and consistent latency. It is specifically engineered to serve as a reliable workhorse for developers who need robust code generation and repository-level context handling without the hardware overhead associated with larger, proprietary models.
The model is built to support agentic workflows and real-world programming tasks, offering a balance between performance and deployment flexibility. Because of its dense architecture, it is well-suited for quantization, making it a practical choice for teams looking to run local AI coding assistants on consumer hardware. With a 128,000-token context window, it effectively manages large codebases, documentation, and complex stack traces, providing a cost-effective solution for developers who require frontier-level coding chops in resource-constrained environments.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Cortecs
The 30B model achieves 59.2% on SWE-bench Verified while running at 80+ tokens per second on MacBooks.
This exact model name is also listed by 17 other providers.