TensorX
An NVIDIA developer-forum thread posted September 8, 2026 by community member azampatti documents an INT4 AutoRound quantization recipe for Qwen3.8-Flash-Next targeting DGX Spark / GB10 hardware. The community fork halves the number of routed experts per token from 10 to 5 and heals the resulting quality loss with a 37 The reported throughput is 60–70 tokens per second in coding workloads, with vLLM launching a GPU KV cache of approximately 644,732 tokens and a maximum concurrency of about 2.46× at 262,144 tokens per request. The variant name Qwen3.8-Flash-Next-125B-A5B-INT4-AutoRound indicates this is a community-modified derivative