Merge Gateway
An NVIDIA developer forum benchmark thread dated August 12, 2026 reports a measured throughput of 116.84 tokens per second for text generation using nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 on an NVIDIA DGX Spark system running vLLM. The post links to the full Spark Arena benchmark entry for the same NVFP4 we Beyond the headline throughput number, the excerpt is light on architectural or methodological detail, presenting a single data point rather than a controlled comparison across configurations or hardware targets. It nevertheless directly names the NVFP4 variant and ties it to a specific deployment stack (vLLM on DGX Sp