Model details
Qwen3 Next 80B A3B Instruct
The Qwen3-Next-80B-A3B-Instruct model represents a significant shift in architectural design, prioritizing scaling efficiency for both training and inference. By utilizing a highly sparse Mixture-of-Experts structure, the model maintains 80 billion total parameters while activating only 3 billion per inference step, resulting in an extremely low activation ratio. This design is complemented by a hybrid attention mechanism that combines Gated DeltaNet and Gated Attention, specifically engineered to handle ultra-long context lengths. These innovations allow the model to deliver high-performance text generation while drastically reducing the computational cost and latency typically associated with large-scale language models.
The development of this model involved rigorous stability optimizations, including zero-centered and weight-decayed layernorm techniques, which were essential for overcoming challenges in reinforcement learning and post-training. The integration of a multi-token prediction mechanism further enhances its performance, accelerating inference speeds and improving overall model utility. With its ability to process up to 256K tokens, the model is particularly well-suited for demanding enterprise applications such as complex document analysis, multi-turn dialogue, and code generation, offering a balance of flagship-level capability and high-throughput production efficiency.
Quick Info
Powered by- Provider
- Qiniu
- Model key
- qwen3-next-80b-a3b-instruct
- Release date
- Sep 12, 2025
- Last updated
- Sep 12, 2025
- Input modalities
- Output modalities
- Capabilities
Limits
- Output tokens
- 32,768 tokens
- Context window
- 131,072 tokens
Latest news about Qwen3 Next 80B A3B Instruct
No articles yet. Fetch the latest news to show it here.