Qwen3.8 2.4T A95B is an open-weight mixture-of-experts model released by Alibaba with 2.4 trillion total parameters and 95 billion activated per token. Its architecture pairs full attention layers with linear-attention layers, a fine-grained expert design that keeps compute and memory bounded as context scales, and a window reaching up to one million tokens. Built-in configurable reasoning controls let developers tune inference depth per request, making the model well suited to coding, large-scale document analysis, and long-running agentic workflows where context grows with tool outputs, retrieved passages, and multi-step traces.
Out of the box on NVIDIA GB300 NVL72 in FP8 precision, the model delivers over 4,000 tokens per second per GPU and over 350 tokens per second per user, with NVFP4 expected to push throughput further. The ecosystem around it is broad: Hugging Face hosts the official weights, NVIDIA NeMo AutoModel enables post-training via full supervised fine-tuning or memory-efficient LoRA on those checkpoints, and open-source inference recipes ship for SGLang, vLLM, and NVIDIA Dynamo, with a model-free NIM container available from NGC. The combination of a trillion-parameter sparse design, hybrid attention for long context, and a mature serving and fine-tuning stack positions the model as a flexible foundation for research and production agentic systems.