Vercel AI Gateway
Alibaba's new coding model Qwen3-Coder-Next achieves the performance of significantly larger models with only 3 billion active parameters.
Model details
Qwen3-Coder-Next is a sparse Mixture-of-Experts model built on Qwen3-Next that carries 80 billion total parameters but activates only 3 billion during inference. Its hybrid architecture combines Gated Attention with Gated DeltaNet, allowing the model to match the coding performance of models requiring 10 to 20 times more active parameters. The design prioritizes coding agent workflows, with a 256,000-token context window, advanced long-horizon reasoning, and specialized training for navigating complex tool usage and recovering from execution failures. This combination of lightweight activation with heavy-weight capability makes it practical for agentic coding tasks in real IDE and CLI environments.
The model was trained through scaled agentic training signals, learning from large collections of verifiable coding tasks paired with executable environments that provide direct feedback. This agent-centric approach includes continued pretraining on code-focused data, supervised fine-tuning on high-quality agent trajectories, and domain-specialized expert training. Qwen3-Coder-Next achieved 70.6% on SWE-Bench and 44.3% on SWE-Bench Pro, competitive performance for its active parameter footprint. Released under an Apache 2.0 license, the open weights are available on Hugging Face in multiple variants and can run locally on consumer hardware, making elite coding agent performance accessible without expensive API costs.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Vercel AI Gateway
Alibaba's new coding model Qwen3-Coder-Next achieves the performance of significantly larger models with only 3 billion active parameters.
Vercel AI Gateway
Alibaba has released Qwen3-Coder-Next, an open-source 80B-parameter coding model that activates just 3B parameters per query, scoring 70.6% on SWE-Bench.
Vercel AI Gateway
The official Qwen Hugging Face model card for Qwen3-Coder-Next-Base confirms the model as an open-weight causal language model designed for coding agents and local development, released by the Qwen team. The card specifies 80B total parameters with 3B activated, 79B non-embedding parameters, a hidden dimension of 2048, Additional architectural details from the model card list 512 experts with 10 activated and 1 shared expert per token, an expert intermediate dimension of 512, and a native context length of 262,144 tokens supporting 370+ programming languages. The card states the model operates in non-thinking mode only and does not e