Model details
Qwen3 Next 80B A3B Thinking
Qwen3 Next 80B A3B Thinking is a reasoning-first model built on the innovative Qwen3-Next architecture, which prioritizes extreme efficiency in both training and inference. The model utilizes a hybrid attention mechanism that combines Gated DeltaNet and Gated Attention to manage long-context modeling effectively. By employing a high-sparsity Mixture-of-Experts structure, it achieves a low activation ratio, requiring only 3 billion parameters during inference despite its 80-billion-parameter scale. This design, complemented by multi-token prediction, allows the model to deliver high throughput and rapid generation, making it well-suited for demanding applications like code synthesis, mathematical proofs, and complex logic.
The model benefits from advanced stability optimizations, including zero-centered and weight-decayed layernorm, which ensure robust performance during pre-training and post-training. Through the application of GSPO, the developers successfully addressed the stability challenges inherent in combining hybrid attention with high-sparsity MoE architectures during reinforcement learning. As a reasoning-enhanced variant, it is specifically tuned to output structured thinking traces, providing a reliable foundation for agentic frameworks, tool use, and retrieval-heavy workflows that require detailed, step-by-step reasoning while maintaining instruction adherence and reducing off-task behavior.
Quick Info
Powered by- Provider
- NovitaAI
- Model key
- qwen/qwen3-next-80b-a3b-thinking
- Release date
- Sep 10, 2025
- Last updated
- Sep 10, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.15
- Output token cost
- $1.50
Limits
- Output tokens
- 32,768 tokens
- Context window
- 131,072 tokens
Latest news about Qwen3 Next 80B A3B Thinking
No articles yet. Fetch the latest news to show it here.