Currently listed through these providers:
Model details
Qwen3-Next 80B-A3B (Thinking)
Qwen3-Next 80B-A3B Thinking marks the debut of the Qwen3-Next series, built as a Mixture-of-Experts language model with 80 billion total parameters and 3 billion activated parameters per forward pass. Its defining architectural choices include a hybrid attention mechanism that combines Gated DeltaNet with Gated Attention, enabling efficient modeling of ultra-long contexts up to roughly 262K tokens. The model also incorporates Multi-Token Prediction during pretraining, which accelerates inference by predicting multiple tokens simultaneously rather than relying on sequential generation. By using an extremely sparse MoE configuration, the design keeps FLOPs per token low while preserving the full capacity a much larger dense model would offer.
The architecture benefits from stability optimizations such as zero-centered and weight-decayed layer normalization, which support robust pretraining and post-training regimes. The model is released with open weights, making it available for both commercial and non-commercial deployments, and carries enhanced reasoning capabilities suited for complex, multi-step tasks. With its combination of long-context capacity, MoE efficiency, and multi-token prediction, this model targets applications in agentic AI where reasoning depth and context handling are both critical.
Quick Info
Powered by- Provider
- Alibaba
- Model key
- qwen3-next-80b-a3b-thinking
- Release date
- Sep 1, 2025
- Last updated
- Sep 1, 2025
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.50
- Output token cost
- $6.00
Limits
- Output tokens
- 32,768 tokens
- Context window
- 131,072 tokens
Transparent token rates
Compare Qwen3-Next 80B-A3B (Thinking) pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Qwen3-Next 80B-A3B (Thinking)
No articles yet. Fetch the latest news to show it here.