Alibaba
Qwen3-Next-80B-A3B-Instruct is Alibaba's latest open-source Mixture-of-Experts (MoE) model, released on September 11, 2025. Despite having 80 billion total
Model details
Qwen3-Next 80B A3B Instruct is an instruction-tuned chat and agent model built on a sparse Mixture-of-Experts architecture, where roughly 3 billion parameters are active per token out of 80 billion total. This low activation ratio lets the model keep high capacity while sharply lowering compute per token, and a hybrid attention design backed by Multi-Token Prediction helps sustain throughput on long inputs. Public reference deployments describe context handling well beyond standard chat windows, with Fireworks noting support up to 262K tokens, making the model a fit for document analysis, repository-scale reasoning, and multi-turn agent workflows that need to keep large working memories in view.
Released in September 2025, the weights are openly available on Hugging Face under an Apache 2.0-aligned license, and the model has been picked up quickly across many inference clouds, including Alibaba Cloud, Fireworks, NVIDIA NIM, DeepInfra, Google Vertex, GMI, Novita, and Parasail. Independent benchmarking across these providers shows consistently strong output speeds in roughly the 180–186 tokens-per-second range at the top end and blended pricing between about $0.18 and $0.29 per million tokens depending on host, so teams can shop for a cost-and-latency profile rather than committing to a single stack. The combination of open weights, agent-friendly context length, and competitive hosted economics positions it as a practical backbone for production assistants, retrieval-heavy tools, and lightweight agent pipelines that want MoE efficiency without sacrificing a long context window.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Alibaba
Qwen3-Next-80B-A3B-Instruct is Alibaba's latest open-source Mixture-of-Experts (MoE) model, released on September 11, 2025. Despite having 80 billion total
Alibaba
Artificial Analysis provides independent third-party benchmarking of Qwen3 Next 80B A3B Instruct across 7 API providers (GMI, Novita, Alibaba Cloud, DeepInfra, Google Vertex, Parasail, and one other), released September 2025. Top output speeds are GMI at 186.1 t/s, Novita at 181.9 t/s, and Alibaba Cloud at 180.4 t/s, w On pricing, the lowest blended price per 1M tokens (7:2:1 cache-input-output) is Parasail at $0.18, followed by DeepInfra at $0.19, with Google Vertex and Alibaba Cloud tied at $0.26, and GMI at $0.29. The page also breaks down cache-hit pricing separately and supports a general agentic 7:2:1 ratio. This is useful, non