Alibaba
Millstone AI published a third-party inference benchmark page for Qwen3-Coder-30B-A3B-Instruct in FP8 precision. The page documents the model as an FP8-quantized 30.5B-parameter Mixture-of-Experts (MoE) architecture built on the Qwen3 base, with 128 experts (8 active per forward pass) and only 3.3B parameters activated The benchmark pages report hardware-specific performance for the FP8 variant, including peak throughput of 334 tok/s on a 1x RTX Pro 6000 Blackwell (96GB), 584 tok/s on a 1x H100 SXM (80GB), and 600 tok/s on a 1x H200 SXM (141GB), tested across concurrency ranges of 1–4 to 1–6 and context lengths from 1K up to 256K. Pe