RunInfra
SGLang and Miles added full Day-0 support for Qwen3.8-2.4T-A95B, Qwen's largest open-source model with 2.4T total parameters and 95B active per token. The post describes a hybrid attention architecture of 92 layers (69 GDN linear-attention layers interleaved with 23 GQA full-attention layers in a 3:1 pattern) and MoE l An NVFP4 checkpoint named RadixArk/Qwen3.8-2.4T-A95B-NVFP4 was released Day-0 alongside the serving stack. Performance work includes FlashInfer kernels (MoE finalize fused with all-reduce and RMSNorm for 10%+ end-to-end gains), a context-parallel GDN prefill kernel, and a low-latency single-GEMM path (4% end-to-end). A
