Currently listed through these providers:
Model details
Qwen3-Next 80B-A3B Instruct
Qwen3-Next 80B-A3B Instruct represents a significant shift in model architecture, focusing on extreme scaling efficiency through a high-sparsity Mixture-of-Experts design. By activating only 3 billion of its 80 billion total parameters per inference step, the model achieves a remarkably low activation ratio of 3.75%, which drastically reduces computational requirements while maintaining high capacity. This architecture is further enhanced by Hybrid Attention, which combines Gated DeltaNet and Gated Attention to manage ultra-long context windows effectively. These design choices make the model particularly well-suited for tasks requiring deep analysis of lengthy documents, complex multi-turn dialogues, and high-throughput code generation.
The development of this model incorporates advanced stability optimizations, including zero-centered and weight-decayed layernorm, to ensure robust performance during both pre-training and post-training phases. The integration of Multi-Token Prediction further accelerates inference speeds and boosts overall performance on downstream tasks. By balancing a massive parameter count with a lightweight active footprint, the model delivers performance comparable to much larger dense counterparts while remaining accessible for cost-conscious enterprise applications. Its ability to handle extensive context lengths with high inference throughput positions it as a versatile tool for production environments that demand both speed and deep analytical capability.
Quick Info
Powered by- Provider
- DevPass (LLM Gateway)
- Model key
- qwen3-next-80b-a3b-instruct
- Release date
- Sep 1, 2025
- Last updated
- Sep 1, 2025
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.15
- Output token cost
- $1.20
Limits
- Output tokens
- 32,768 tokens
- Context window
- 131,072 tokens
Transparent token rates
Compare Qwen3-Next 80B-A3B Instruct pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Qwen3-Next 80B-A3B Instruct
No articles yet. Fetch the latest news to show it here.