Hetzner
Dre Dyson's hands-on deployment guide focuses on serving Qwen3.6-35B-A3B-FP8 via vLLM, comparing configurations against a Qwen3.5 baseline. The author rebuilt vLLM from dev wheels (0.19.1rc1) using FlashInfer as the attention backend and the Qwen/Qwen3.6-35B-A3B-FP8 checkpoint on a dual-NVIDIA-GPU rig. Peak throughput The article highlights two model-level improvements called out in the Qwen3.6 release notes: agentic coding enhancements for frontend workflows and repository-level reasoning, and a thinking-preservation feature that retains reasoning context from prior conversation turns. It emphasizes the importance of running an est