Currently listed through these providers:
Model details
Qwen3-Next-80B-A3B-Thinking
The Qwen3-Next series introduces a fresh architectural approach built around hybrid attention, which pairs Gated DeltaNet with Gated Attention to model ultra-long contexts more efficiently than standard attention alone. This hybrid mechanism is combined with a highly sparse Mixture-of-Experts structure, enabling the model to draw on an extreme low activation ratio that dramatically cuts FLOPs per token while keeping the full parameter capacity intact. Additional stability measures such as zero-centered and weight-decayed layer normalization help keep training on track despite the unconventional architecture, and a Multi-Token Prediction head accelerates inference by generating multiple tokens per forward pass rather than relying on autoregressive decoding alone.
Built atop this foundation, the Qwen3-Next-80B-A3B-Base checkpoint was post-trained into two distinct variants, with the Thinking variant tailored for extended reasoning tasks. Early benchmarking shows this 80B model activating only about 3B parameters per token—yet delivering performance on par with or slightly above the denser Qwen3-32B, while consuming under ten percent of the training GPU hours. The efficiency gains become even more pronounced at context lengths beyond 32K tokens, where the architecture delivers over ten times the throughput of comparable dense models. The model ships in FP8-quantized form with fine-grained per-block quantization for practical deployment, and its open-weight status invites community fine-tuning and experimentation.
Quick Info
Powered by- Provider
- Hugging Face
- Model key
- Qwen/Qwen3-Next-80B-A3B-Thinking
- Release date
- Sep 11, 2025
- Last updated
- Sep 11, 2025
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.30
- Output token cost
- $2.00
Limits
- Output tokens
- 131,072 tokens
- Context window
- 262,144 tokens
Transparent token rates
Compare Qwen3-Next-80B-A3B-Thinking pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Qwen3-Next-80B-A3B-Thinking
No articles yet. Fetch the latest news to show it here.