QVAC
A practical Medium tutorial demonstrates running Qwen3.6-35B-A3B locally via llama.cpp on consumer hardware with just 6 GB of VRAM and 32 GB of RAM, achieving approximately 30 tokens per second. The article, published May 10, 2026, highlights that recent updates to llama.cpp combined with the model's sparse 35B/3B MoE The piece frames this as a significant improvement over prior setups where larger models required substantial VRAM and were impractical locally, while smaller models were limited to 32K–128K contexts. By leveraging Qwen3.6-35B-A3B's sparsity and recent llama.cpp optimizations, the author shows a path to running frontie