Sulat.com
AI models
Get 10-25% off
Get 10-25% off from Qwen
Alibaba logo

Model details

Qwen3-Next 80B-A3B Instruct

Qwen3-Next 80B A3B Instruct is an instruction-tuned chat and agent model built on a sparse Mixture-of-Experts architecture, where roughly 3 billion parameters are active per token out of 80 billion total. This low activation ratio lets the model keep high capacity while sharply lowering compute per token, and a hybrid attention design backed by Multi-Token Prediction helps sustain throughput on long inputs. Public reference deployments describe context handling well beyond standard chat windows, with Fireworks noting support up to 262K tokens, making the model a fit for document analysis, repository-scale reasoning, and multi-turn agent workflows that need to keep large working memories in view.

Released in September 2025, the weights are openly available on Hugging Face under an Apache 2.0-aligned license, and the model has been picked up quickly across many inference clouds, including Alibaba Cloud, Fireworks, NVIDIA NIM, DeepInfra, Google Vertex, GMI, Novita, and Parasail. Independent benchmarking across these providers shows consistently strong output speeds in roughly the 180–186 tokens-per-second range at the top end and blended pricing between about $0.18 and $0.29 per million tokens depending on host, so teams can shop for a cost-and-latency profile rather than committing to a single stack. The combination of open weights, agent-friendly context length, and competitive hosted economics positions it as a practical backbone for production assistants, retrieval-heavy tools, and lightweight agent pipelines that want MoE efficiency without sacrificing a long context window.

Alibabaqwen3-next-80b-a3b-instructqwen

Quick Info

Powered by
Provider
Alibaba
Model key
qwen3-next-80b-a3b-instruct
Release date
Sep 1, 2025
Last updated
Sep 1, 2025
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.50
Output token cost
$2.00

Limits

Output tokens
32,768 tokens
Context window
131,072 tokens

Transparent token rates

Compare Qwen3-Next 80B-A3B Instruct pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3-Next 80B-A3B Instruct

Alibaba

CoverageComparison

Qwen3-Next-80B-A3B-Instruct is Alibaba's latest open-source Mixture-of-Experts (MoE) model, released on September 11, 2025. Despite having 80 billion total

Alibaba

CoverageBenchmark

Artificial Analysis provides independent third-party benchmarking of Qwen3 Next 80B A3B Instruct across 7 API providers (GMI, Novita, Alibaba Cloud, DeepInfra, Google Vertex, Parasail, and one other), released September 2025. Top output speeds are GMI at 186.1 t/s, Novita at 181.9 t/s, and Alibaba Cloud at 180.4 t/s, w On pricing, the lowest blended price per 1M tokens (7:2:1 cache-input-output) is Parasail at $0.18, followed by DeepInfra at $0.19, with Google Vertex and Alibaba Cloud tied at $0.26, and GMI at $0.29. The page also breaks down cache-hit pricing separately and supports a general agentic 7:2:1 ratio. This is useful, non

Videos about Qwen3-Next 80B-A3B Instruct

More models around Qwen3-Next 80B-A3B Instruct