Currently listed through these providers:
Model details
Qwen3-Next 80B-A3B (Thinking)
Qwen3-Next 80B-A3B Thinking represents a significant shift in model architecture, designed to balance massive parameter capacity with extreme operational efficiency. By utilizing a high-sparsity mixture-of-experts structure, the model activates only 3 billion parameters during inference, which drastically reduces the computational cost per token. This design is further enhanced by a hybrid attention mechanism that replaces standard attention with a combination of Gated DeltaNet and Gated Attention, allowing the model to handle ultra-long context lengths with high throughput. These innovations ensure that the model remains both powerful and agile, making it well-suited for complex tasks that require deep reasoning and extensive information processing.
The development of this model involved rigorous stability optimizations, including zero-centered and weight-decayed layernorm techniques, which were critical for maintaining performance during both pre-training and reinforcement learning phases. By leveraging multi-token prediction, the model achieves faster inference speeds, while the integration of GSPO methods addresses the specific stability challenges inherent in training high-sparsity architectures. These advancements in post-training and architectural design allow the model to deliver performance that rivals dense alternatives while requiring only a fraction of the training resources. As a result, it serves as a robust foundation for demanding agentic applications that require reliable, high-speed reasoning across diverse language tasks.
Quick Info
Powered by- Provider
- DevPass (LLM Gateway)
- Model key
- qwen3-next-80b-a3b-thinking
- Release date
- Sep 1, 2025
- Last updated
- Sep 1, 2025
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.15
- Output token cost
- $1.20
Limits
- Output tokens
- 32,768 tokens
- Context window
- 131,072 tokens
Transparent token rates
Compare Qwen3-Next 80B-A3B (Thinking) pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.