Hugging Face
Alibaba's new coding model Qwen3-Coder-Next achieves the performance of significantly larger models with only 3 billion active parameters.
Model details
Qwen3-Coder-Next is a specialized language model engineered to serve as a high-performance engine for coding agents and local development environments. Built upon a hybrid architecture that integrates Mixture-of-Experts with gated attention and DeltaNet layers, the model manages 80 billion total parameters while activating only 3 billion during each inference pass. This design intent focuses on delivering elite coding capabilities—such as long-horizon reasoning and complex tool usage—within a compact active footprint, allowing developers to run sophisticated agentic workflows on consumer-grade hardware.
The model’s development lineage centers on a rigorous agentic training recipe that emphasizes learning from environment feedback. By utilizing large-scale synthesis of verifiable coding tasks and executable environments, the model underwent continued pre-training and supervised fine-tuning on high-quality agent trajectories. This approach enables the model to excel at recovering from execution failures and navigating dynamic coding tasks. With its support for extensive context lengths and compatibility with various CLI and IDE platforms, it is positioned as a versatile tool for both individual developers and enterprise-scale coding agent deployments.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Hugging Face
Alibaba's new coding model Qwen3-Coder-Next achieves the performance of significantly larger models with only 3 billion active parameters.
Hugging Face
Qwen team has just released Qwen3-Coder-Next, an open-weight language model designed for coding agents and local development. It sits on top of the...
Hugging Face
Artificial Analysis evaluates Qwen3 Coder Next as an open-weights, non-reasoning model from Alibaba released in February 2026, with 79.7 billion total parameters and 3 billion active parameters per token. It supports text input and output with a 256k-token context window and is distributed under the Apache 2.0 license The model scores 10 on the Artificial Analysis Intelligence Index, placing it well above the median among comparable open-weight non-reasoning models, though it is described as slower than average at roughly 102 output tokens per second and notably verbose. A reasoning variant may also exist but is not covered on the s