Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
RunInfra logo

Model details

Qwen3.8 2.4T A95B (NVFP4)

This checkpoint is presented as a massively scaled mixture-of-experts language model, with 1.2608 trillion base parameters in an NVFP4-quantized artifact. Its reported architecture, lineage from the Qwen3.8 2.4T base, ModelOpt quantization, and sharded safetensors format make it better suited to substantial server infrastructure than ordinary desktop inference.

The available evidence supports a practical profile centered on long-context conversational use and large-scale model serving, while the reported 256K context window and roughly 294.4 GB VRAM footprint indicate meaningful deployment requirements. The four-bit representation reduces the storage burden relative to the original large checkpoint, but it remains a multi-terabyte-scale artifact; hosted access or carefully provisioned hardware is the more realistic fit.

RunInfraInferact/Qwen3.8-2.4T-A95B-NVFP4qwen

Quick Info

Powered by
Provider
RunInfra
Model key
Inferact/Qwen3.8-2.4T-A95B-NVFP4
Release date
Aug 12, 2026
Last updated
Aug 12, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$2.00
Output token cost
$6.00

Limits

Output tokens
32,768 tokens
Context window
262,144 tokens

Latest news about Qwen3.8 2.4T A95B (NVFP4)

RunInfra

Coverage

SGLang and Miles added full Day-0 support for Qwen3.8-2.4T-A95B, Qwen's largest open-source model with 2.4T total parameters and 95B active per token. The post describes a hybrid attention architecture of 92 layers (69 GDN linear-attention layers interleaved with 23 GQA full-attention layers in a 3:1 pattern) and MoE l An NVFP4 checkpoint named RadixArk/Qwen3.8-2.4T-A95B-NVFP4 was released Day-0 alongside the serving stack. Performance work includes FlashInfer kernels (MoE finalize fused with all-reduce and RMSNorm for 10%+ end-to-end gains), a context-parallel GDN prefill kernel, and a low-latency single-GEMM path (4% end-to-end). A

RunInfra

Coverage

vLLM announced Day-0 support for Qwen3.8-2.4T-A95B on August 12, 2026, describing it as the first Qwen-Max-class model released as open weights and noting it reuses the Qwen 3.5 architecture so it runs on vLLM out of the box. The model is a 2.4-trillion-parameter sparse mixture-of-experts with 512 experts and a 92-laye Running Qwen3.8-2.4T-A95B at FP8 or BF16 requires at least two NVIDIA B300 or AMD MI355X nodes, while the FP4 quantized version can run on a single node. Inferact quantized selected layers, including routed experts, to FP4 using Round-to-Nearest with activation calibration to enable 4-bit activations. vLLM reported GSM

Videos about Qwen3.8 2.4T A95B (NVFP4)

More models around Qwen3.8 2.4T A95B (NVFP4)