Alibaba (China)
Qwen3-Next-80B-A3B-Instruct is Alibaba's latest open-source Mixture-of-Experts (MoE) model, released on September 11, 2025. Despite having 80 billion total
Model details
The Qwen3-Next 80B-A3B Instruct model represents a significant shift toward scaling efficiency in large language models. Built on a high-sparsity Mixture-of-Experts architecture, it manages 80 billion total parameters while activating only 3 billion per inference step, which drastically reduces computational overhead. The model integrates a hybrid attention mechanism—combining Gated DeltaNet and Gated Attention—to handle ultra-long context windows effectively. This design intent focuses on delivering flagship-level performance for tasks such as long document analysis, code generation, and multi-turn dialogues, ensuring that the system remains both powerful and highly responsive in production environments.
The model benefits from advanced training and post-training optimizations, including multi-token prediction to accelerate inference and stability-focused techniques like zero-centered, weight-decayed layernorm. These enhancements allow the model to maintain robust performance during reinforcement learning and alignment, solving long-standing stability issues associated with sparse architectures. By balancing high-throughput capabilities with the ability to process extensive context, the model is well-suited for agentic workflows, retrieval-augmented generation, and complex instruction-following tasks. It serves as a versatile assistant that provides deterministic, high-quality outputs, making it a practical choice for enterprise applications requiring both speed and deep reasoning.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Alibaba (China)
Qwen3-Next-80B-A3B-Instruct is Alibaba's latest open-source Mixture-of-Experts (MoE) model, released on September 11, 2025. Despite having 80 billion total