LowRouter
An NVIDIA DGX Spark user forum post benchmarks Mistral Medium 3.5 128B NVFP4 with the official EAGLE draft model on a single DGX Spark / GB10 node using spark-vllm-docker TF5 and vLLM 0.20.2rc1. Target model memory was 72.91 GiB, KV cache 89,504 tokens, and max concurrency 5.46x at 16k context. The --load-format auto path was stable at 559 seconds load time, while --load-format fastsafetensors triggered OOM at 16k context. Generation throughput at temperature 0.0 ran about 6.9 tok/s for 256-token Q&A, 8.9 tok/s for code and math, and 9.1 tok/s for 2048-token LongCode, all reproduced across two runs. Prompt processing on 3,026 warmup tokens reached 166.2 tok/s, establishing a lower-bound TTFT benchmark for NVFP4 plus EAGLE speculative decoding on consumer-grade DGX Spark silicon.
