Currently listed through these providers:
Model details
Qwen3 Next 80B A3B Instruct
As the first installment in the Qwen3-Next series, Qwen3-Next-80B-A3B-Instruct introduces a next-generation architecture aimed at scaling efficiently to longer contexts and more agentic workloads. The model replaces standard attention with a hybrid combination of Gated DeltaNet and Gated Attention, designed for efficient ultra-long-context modeling, and layers on a high-sparsity Mixture-of-Experts design that keeps the active parameter count low per token while preserving overall model capacity. Stability enhancements such as zero-centered and weight-decayed layernorm support more robust pre-training and post-training, and multi-token prediction is included to improve pretraining performance and speed up inference. Together these choices mark a clear architectural step beyond the earlier Qwen3 dense line, focusing on inference efficiency at very long context lengths rather than pure parameter count.
In practice, the model is positioned as a long-context instruction-tuned LLM that delivers quality comparable to much larger Qwen3 instruct variants while being lighter to run. According to the published highlights, Qwen3-Next-80B-A3B-Instruct performs on par with Qwen3-235B-A22B-Instruct-2507 on certain benchmarks and shows notable advantages on ultra-long-context tasks up to 256K tokens. The base variant is reported to outperform Qwen3-32B-Base on downstream tasks using roughly 10% of the total training cost while delivering about ten times the inference throughput at context lengths above 32K. As an open-weight release with an NVIDIA NGC container deployment path, it is a practical fit for teams that need strong long-context instruction following and agentic behavior without the compute footprint of the largest flagship models.
Quick Info
Powered by- Provider
- Jiekou.AI
- Model key
- qwen/qwen3-next-80b-a3b-instruct
- Release date
- Jan 1, 2026
- Last updated
- Jan 1, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.15
- Output token cost
- $1.50
Limits
- Output tokens
- 65,536 tokens
- Context window
- 65,536 tokens