Currently listed through these providers:
Model details
Qwen3 Next 80B A3B Thinking
Qwen3 Next 80B A3B Thinking is a specialized reasoning model built on the innovative Qwen3-Next architecture, which prioritizes extreme efficiency in both training and inference. The model employs a hybrid attention mechanism that combines Gated DeltaNet and Gated Attention to manage long-context modeling effectively. By utilizing a high-sparsity Mixture-of-Experts structure, it achieves a low activation ratio, allowing the 80-billion-parameter model to activate only 3 billion parameters during inference. This design, complemented by multi-token prediction to accelerate generation, makes it particularly well-suited for demanding tasks such as mathematical proofs, code synthesis, and complex agentic planning.
The model benefits from advanced stability optimizations, including zero-centered and weight-decayed layernorm, which ensure robust performance during pre-training and post-training phases. By leveraging GSPO, the development team successfully addressed the stability challenges inherent in combining hybrid attention with high-sparsity MoE during reinforcement learning. As a reasoning-first chat model, it is tuned to output structured thinking traces, providing a reliable framework for retrieval-heavy workflows and multi-step logic. Its ability to maintain stability across long chains of thought while reducing off-task behavior positions it as a powerful tool for developers building sophisticated agent frameworks and standardized benchmarking applications.
Quick Info
Powered by- Provider
- Vercel AI Gateway
- Model key
- alibaba/qwen3-next-80b-a3b-thinking
- Release date
- Sep 1, 2025
- Last updated
- Sep 1, 2025
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.15
- Output token cost
- $1.20
Limits
- Output tokens
- 262,144 tokens
- Context window
- 262,144 tokens
Transparent token rates
Compare Qwen3 Next 80B A3B Thinking pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Qwen3 Next 80B A3B Thinking
No articles yet. Fetch the latest news to show it here.