LMStudio
Millstone AI published a detailed inference benchmark for Qwen3-Coder-30B-A3B-Instruct in FP8 precision, documenting the model's architecture and hardware performance. The page describes it as a 30.5B-parameter Mixture-of-Experts model with 128 experts (8 active per forward pass) and 3.3B activated parameters at runtim The benchmark reports peak throughput of 334 tok/s on a single RTX Pro 6000 Blackwell 96GB, 584 tok/s on a single H100 SXM 80GB, and 600 tok/s on a single H200 SXM 141GB, with tested concurrency up to 6 and context lengths ranging from 1K up to 256K. Capacity-planning tables indicate practical concurrent request counts