Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
OpenRouter logo

Model details

Qwen3-Next 80B-A3B (Thinking)

Qwen3-Next 80B-A3B Thinking belongs to Alibaba's Qwen3 family of large language models, sharing the underlying Qwen3-Next architecture that a third-party host describes as the first generation built on that innovative design. The architecture combines a hybrid attention system in a roughly 3:1 ratio, pairing Gated DeltaNet linear attention for efficient long-sequence processing with standard Gated Attention for accurate information retrieval. To scale efficiently, the model uses an ultra-sparse Mixture-of-Experts layout with 512 experts and a small routed subset active per token, letting it deliver high capacity while keeping inference compute relatively low for its parameter class.

As the Thinking-tuned sibling in the Qwen3-Next 80B-A3B line, this variant is positioned for step-by-step reasoning workloads rather than straightforward instruction following, making it well suited to complex analytical tasks, multi-step problem solving, and tool-augmented workflows that benefit from explicit chain-of-thought behavior. Its open weights, supported context window, and text-in/text-out interface make it practical for self-hosted deployments, while the hybrid attention design aims to balance long-context efficiency with retrieval accuracy—an area where pure linear attention traditionally struggles. Developers integrating this model can expect a reasoning-first profile that trades some raw throughput for stronger deliberation on harder prompts, fitting naturally into agent pipelines that mix structured output, temperature control, and external tool calls.

OpenRouterqwen/qwen3-next-80b-a3b-thinkingqwen

Quick Info

Powered by
Provider
OpenRouter
Model key
qwen/qwen3-next-80b-a3b-thinking
Release date
Sep 1, 2025
Last updated
Sep 1, 2025
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$1.20

Limits

Output tokens
235,929 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen3-Next 80B-A3B (Thinking) pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3-Next 80B-A3B (Thinking)

Kilo Gateway

CoverageBenchmark

The Together AI model card for Qwen3-Next-80B-A3B-Thinking documents the exact variant as a next-generation reasoning model with extreme efficiency, explicitly supporting thinking mode only with automatic tag inclusion. Architecture details include 48 layers, a 2048 hidden dimension with a hybrid layout pattern, 512 to The card lists 262K native context length extensible to 1M tokens via YaRN scaling, more than 10x higher throughput on contexts over 32K tokens, and deployment support via SGLang and vLLM with Multi-Token Prediction. Model provider is listed as Qwen. Targeted applications include scientific research, complex data analy

Videos about Qwen3-Next 80B-A3B (Thinking)

More models around Qwen3-Next 80B-A3B (Thinking)