Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Jalapeno Cloud logo

Model details

Qwen3-Next 80B-A3B (Thinking)

Built on the same highly sparse mixture-of-experts backbone as its Instruct sibling, this variant is post-trained specifically for complex reasoning chains and operates exclusively in thinking mode, automatically wrapping its analytical output in the corresponding tags. Its architecture uses a 48-layer hybrid layout with a 2048 hidden dimension and a multi-token prediction mechanism that accelerates token generation, while 512 total experts with only 10 plus one shared activated per layer keep per-token compute very low. The same recipe draws on roughly 15 trillion tokens of specialized reasoning-focused post-training, allowing the model to produce longer, more deliberate thinking traces than earlier Qwen3 generations without sacrificing throughput.

In comparative testing, this thinking-tuned release leads the Instruct counterpart on 18 of 21 shared benchmarks, including AIME 2025, GPQA, HMMT25, MMLU-Pro, LiveCodeBench v6, and the Tau2 and TAU-bench agent suites, while the Instruct variant only edges ahead on writing-style tasks such as Arena-Hard v2 and WritingBench. Native context reaches 262K tokens and extends toward 1M with YaRN scaling, and the model advertises more than ten times higher throughput than predecessors once inputs cross 32K tokens, making it well suited for deep research, multi-step agent workflows, and long-document analysis. Open weights and SGLang or vLLM deployment support give infrastructure teams flexibility to self-host, while the open reasoning focus makes the model a practical choice when traceable, step-by-step analytical output matters more than conversational polish.

Jalapeno CloudQwen3-Next-80B-A3B-Thinkingqwen

Quick Info

Powered by
Provider
Jalapeno Cloud
Model key
Qwen3-Next-80B-A3B-Thinking
Release date
Sep 1, 2025
Last updated
Sep 1, 2025
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$1.50

Limits

Output tokens
32,768 tokens
Context window
131,072 tokens

Transparent token rates

Compare Qwen3-Next 80B-A3B (Thinking) pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3-Next 80B-A3B (Thinking)

No articles yet. Fetch the latest news to show it here.

Videos about Qwen3-Next 80B-A3B (Thinking)

More models around Qwen3-Next 80B-A3B (Thinking)