Hugging Face
Qwen3-Next-80B-A3B-Instruct is Alibaba's latest open-source Mixture-of-Experts (MoE) model, released on September 11, 2025. Despite having 80 billion total
Model details
Qwen3-Next-80B-A3B-Instruct represents a shift toward extreme scaling efficiency, utilizing a specialized architecture designed to handle massive context windows while minimizing computational overhead. At its core, the model employs a high-sparsity Mixture-of-Experts structure that activates only 3 billion of its 80 billion total parameters during any single inference step. This design is bolstered by a hybrid attention mechanism, which combines Gated DeltaNet and Gated Attention to manage long-range dependencies effectively. By integrating multi-token prediction, the model achieves significant gains in inference speed, making it well-suited for high-throughput production environments that require both depth and rapid response times.
The development of this model focused on overcoming traditional bottlenecks in training and inference stability, particularly when scaling to long-context tasks. Through the implementation of stability-focused techniques like zero-centered and weight-decayed layer normalization, the model maintains robust performance during both pre-training and post-training phases. These advancements allow the model to deliver performance comparable to much larger dense models while consuming a fraction of the training cost. Its ability to process extensive inputs makes it a practical choice for complex document analysis, multi-turn dialogue, and code generation, offering a balance of high-level reasoning capabilities and operational economy.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Hugging Face
Qwen3-Next-80B-A3B-Instruct is Alibaba's latest open-source Mixture-of-Experts (MoE) model, released on September 11, 2025. Despite having 80 billion total