Venice AI
Alibaba released the open weights for Qwen3.8-2.4T-A95B, the open-weight variant of Qwen3.8 Max and a 2.4 trillion-parameter mixture-of-experts model with 95 billion activated parameters per token, as described in an NVIDIA Technical Blog post dated August 12, 2026. The architecture combines fine-grained MoE routing with a hybrid full and linear attention design, supports a context window of up to one million tokens, and produces outputs of up to 128K tokens for reasoning and agentic workloads. The same NVIDIA Technical Blog explains that, without additional tuning, Qwen3.8-2.4T-A95B achieves over 4,000 tokens per second per GPU and more than 350 tokens per second per user on NVIDIA GB300 NVL72 systems in FP8 precision on Day 0, with further gains expected from NVFP4 optimizations. NVIDIA NeMo AutoModel supports post-training via full supervised fine-tuning or memory-efficient LoRA, and open-source inference recipes are available for SGLang, vLLM, and NVIDIA Dynamo, with additional deployment via a model-free NVIDIA NIM container from NVIDIA NGC.