DevPass (LLM Gateway)
Community users on the NVIDIA DGX Spark forum report running Qwen3.5-122B-A10B on a single Spark node, achieving up to 51 tokens per second with vLLM patches and custom quick-start configurations. The deployment thread documents hands-on local inference performance for the MoE model on GB10 hardware. Thread participants compare vLLM 0.19.1-stable against 0.20.1 builds specifically for this model. Testing showed approximately a 9% performance regression when moving from vLLM 0.19.1 to 0.20.1 on Qwen3.5-122B-A10B, with Q&A dropping 8.3%, code 9.6%, and long-context 10.5%. Users noted vLLM 0.20.1 stabilized tool calls for stock templates but degraded MTP-enabled tool quality. Maintainers recommended staying on 0.19.1-stable pending further regression benchmarks.