Currently listed through these providers:
Model details
Qwen3-Next 80B-A3B Instruct
The Qwen3-Next series was introduced by the Alibaba Qwen team as a next-generation effort focused on scaling efficiency, with the 80B-A3B Instruct release serving as the first installment. It replaces standard attention with a hybrid design that combines Gated DeltaNet and Gated Attention, allowing the model to model long contexts more efficiently. It also relies on a high-sparsity Mixture-of-Experts layout that drives a very low activation ratio per token, preserving overall capacity while cutting compute. Stability techniques such as zero-centered and weight-decayed layernorm, along with multi-token prediction, help both pretraining and post-training behave predictably, and the Instruct checkpoint is tuned to behave well on agentic and ultra-long context workloads up to 256K tokens while matching much larger dense-style Qwen3 chat models on selected benchmarks.
In practical terms, the model pairs this efficient long-context architecture with a predictable deployment profile, making it a sensible fit for applications that need to reason over long documents, codebases, or tool-augmented agent loops without paying the full cost of the largest Qwen3 chat variants. Its open-weights nature makes it attractive to teams that want to self-host, fine-tune, or audit the behavior of a frontier-style instruct model, and it can be paired with external tools for retrieval, structured output, and step-by-step planning. Buyers evaluating it against newer dense alternatives should weigh context handling, throughput on long inputs, and total cost rather than headline parameter count alone.
Quick Info
Powered by- Provider
- Jalapeno Cloud
- Model key
- Qwen3-Next-80B-A3B-Instruct
- Release date
- Sep 1, 2025
- Last updated
- Sep 1, 2025
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.15
- Output token cost
- $1.50
Limits
- Output tokens
- 32,768 tokens
- Context window
- 129,024 tokens
Transparent token rates
Compare Qwen3-Next 80B-A3B Instruct pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Qwen3-Next 80B-A3B Instruct
No articles yet. Fetch the latest news to show it here.
