Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Vercel AI Gateway logo

Model details

Qwen3 Next 80B A3B Thinking

Qwen3 Next 80B A3B Thinking is a specialized reasoning model built on the innovative Qwen3-Next architecture, which prioritizes extreme efficiency in both training and inference. The model employs a hybrid attention mechanism that combines Gated DeltaNet and Gated Attention to manage long-context modeling effectively. By utilizing a high-sparsity Mixture-of-Experts structure, it achieves a low activation ratio, allowing the 80-billion-parameter model to activate only 3 billion parameters during inference. This design, complemented by multi-token prediction to accelerate generation, makes it particularly well-suited for demanding tasks such as mathematical proofs, code synthesis, and complex agentic planning.

The model benefits from advanced stability optimizations, including zero-centered and weight-decayed layernorm, which ensure robust performance during pre-training and post-training phases. By leveraging GSPO, the development team successfully addressed the stability challenges inherent in combining hybrid attention with high-sparsity MoE during reinforcement learning. As a reasoning-first chat model, it is tuned to output structured thinking traces, providing a reliable framework for retrieval-heavy workflows and multi-step logic. Its ability to maintain stability across long chains of thought while reducing off-task behavior positions it as a powerful tool for developers building sophisticated agent frameworks and standardized benchmarking applications.

Vercel AI Gatewayalibaba/qwen3-next-80b-a3b-thinkingqwen

Quick Info

Powered by
Provider
Vercel AI Gateway
Model key
alibaba/qwen3-next-80b-a3b-thinking
Release date
Sep 1, 2025
Last updated
Sep 1, 2025
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$1.20

Limits

Output tokens
262,144 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen3 Next 80B A3B Thinking pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3 Next 80B A3B Thinking

No articles yet. Fetch the latest news to show it here.

Videos about Qwen3 Next 80B A3B Thinking

More models around Qwen3 Next 80B A3B Thinking