Currently listed through these providers:
Model details
Qwen3-Next 80B-A3B (Thinking)
Built on the same highly sparse mixture-of-experts backbone as its Instruct sibling, this variant is post-trained specifically for complex reasoning chains and operates exclusively in thinking mode, automatically wrapping its analytical output in the corresponding tags. Its architecture uses a 48-layer hybrid layout with a 2048 hidden dimension and a multi-token prediction mechanism that accelerates token generation, while 512 total experts with only 10 plus one shared activated per layer keep per-token compute very low. The same recipe draws on roughly 15 trillion tokens of specialized reasoning-focused post-training, allowing the model to produce longer, more deliberate thinking traces than earlier Qwen3 generations without sacrificing throughput.
In comparative testing, this thinking-tuned release leads the Instruct counterpart on 18 of 21 shared benchmarks, including AIME 2025, GPQA, HMMT25, MMLU-Pro, LiveCodeBench v6, and the Tau2 and TAU-bench agent suites, while the Instruct variant only edges ahead on writing-style tasks such as Arena-Hard v2 and WritingBench. Native context reaches 262K tokens and extends toward 1M with YaRN scaling, and the model advertises more than ten times higher throughput than predecessors once inputs cross 32K tokens, making it well suited for deep research, multi-step agent workflows, and long-document analysis. Open weights and SGLang or vLLM deployment support give infrastructure teams flexibility to self-host, while the open reasoning focus makes the model a practical choice when traceable, step-by-step analytical output matters more than conversational polish.
Quick Info
Powered by- Provider
- Jalapeno Cloud
- Model key
- Qwen3-Next-80B-A3B-Thinking
- Release date
- Sep 1, 2025
- Last updated
- Sep 1, 2025
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.15
- Output token cost
- $1.50
Limits
- Output tokens
- 32,768 tokens
- Context window
- 131,072 tokens
Transparent token rates
Compare Qwen3-Next 80B-A3B (Thinking) pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Qwen3-Next 80B-A3B (Thinking)
No articles yet. Fetch the latest news to show it here.
