Jalapeno Cloud
The Millstone AI inference benchmark page documents Qwen3.5-35B-A3B-FP8 as a 35B-parameter Mixture-of-Experts model with 256 experts (8 routed plus 1 shared active per forward pass) and 3B activated parameters, organized across 40 transformer layers. It confirms the hybrid Gated Delta Networks architecture combined wit The page positions Qwen3.5-35B-A3B as competitive with much larger models such as Qwen3-235B-A22B across math, coding, and multilingual benchmarks, and provides measured throughput figures across hardware configurations including 1x and 2x RTX Pro 6000 Blackwell, 1x H100 SXM, and 1x H200 SXM, with peak throughput reach
