Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Hugging Face logo

Model details

Qwen3-Next-80B-A3B-Thinking

The Qwen3-Next series introduces a fresh architectural approach built around hybrid attention, which pairs Gated DeltaNet with Gated Attention to model ultra-long contexts more efficiently than standard attention alone. This hybrid mechanism is combined with a highly sparse Mixture-of-Experts structure, enabling the model to draw on an extreme low activation ratio that dramatically cuts FLOPs per token while keeping the full parameter capacity intact. Additional stability measures such as zero-centered and weight-decayed layer normalization help keep training on track despite the unconventional architecture, and a Multi-Token Prediction head accelerates inference by generating multiple tokens per forward pass rather than relying on autoregressive decoding alone.

Built atop this foundation, the Qwen3-Next-80B-A3B-Base checkpoint was post-trained into two distinct variants, with the Thinking variant tailored for extended reasoning tasks. Early benchmarking shows this 80B model activating only about 3B parameters per token—yet delivering performance on par with or slightly above the denser Qwen3-32B, while consuming under ten percent of the training GPU hours. The efficiency gains become even more pronounced at context lengths beyond 32K tokens, where the architecture delivers over ten times the throughput of comparable dense models. The model ships in FP8-quantized form with fine-grained per-block quantization for practical deployment, and its open-weight status invites community fine-tuning and experimentation.

Hugging FaceQwen/Qwen3-Next-80B-A3B-Thinkingqwen

Quick Info

Powered by
Provider
Hugging Face
Model key
Qwen/Qwen3-Next-80B-A3B-Thinking
Release date
Sep 11, 2025
Last updated
Sep 11, 2025
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.30
Output token cost
$2.00

Limits

Output tokens
131,072 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen3-Next-80B-A3B-Thinking pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3-Next-80B-A3B-Thinking

No articles yet. Fetch the latest news to show it here.

Videos about Qwen3-Next-80B-A3B-Thinking

More models around Qwen3-Next-80B-A3B-Thinking