Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Cortecs logo

Model details

Qwen3 Coder Next

Qwen3 Coder Next 80B is a Mixture of Experts coding model that carries 80 billion total parameters but activates only 3 billion per forward pass, enabling it to match the coding output of much larger dense models while running on hardware accessible to individual developers. The model was designed from the ground up for coding agents and local development workflows, with a 256,000-token context window that lets it reason across large codebases and extended conversations without losing thread. Its architecture supports tool calling and structured JSON output, making it suitable for integrating into autonomous coding pipelines where the model must plan, execute, and verify actions across multiple steps.

The model was trained on over 7.5 trillion tokens with more than 70% code-focused data, giving it deep exposure to software engineering patterns, repository structures, and debugging scenarios. It achieves competitive scores on real-world software engineering benchmarks, reaching performance levels that rival models with significantly higher activated parameter counts. Being fully open-weight under the Apache 2.0 license, it can be quantized and deployed locally on consumer GPUs or developer workstations, eliminating reliance on external API calls for teams that need data privacy or cost control. This combination of benchmark-competitive coding ability, long-horizon context, and self-hostable design positions it as a practical choice for developers building autonomous agents, CI/CD automation, or offline coding assistants.

Cortecsqwen3-coder-nextqwen

Quick Info

Powered by
Provider
Cortecs
Model key
qwen3-coder-next
Release date
Feb 3, 2026
Last updated
Feb 3, 2026
Knowledge cutoff
2025-09
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.167
Output token cost
$0.891

Limits

Output tokens
256,000 tokens
Context window
256,000 tokens

Transparent token rates

Compare Qwen3 Coder Next pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3 Coder Next

Cortecs

Coverage

The Qwen team's official Hugging Face model card for Qwen3-Coder-Next-Base announces the model as an open-weight causal language model purpose-built for coding agents and local development. It pairs a sparse Mixture-of-Experts layout (512 experts, 10 activated plus 1 shared) with a hybrid attention scheme that interlea The same card emphasizes agentic coding strengths enabled by training on diverse, broad-coverage corpora: robust tool calling, scaffold and template adaptation, and error detection and recovery, all intended to serve as a backbone for reliable coding agents. It notes the model operates in non-thinking mode only and rec

Cortecs

CoverageBenchmark

OpenRouter's listing confirms Qwen3 Coder Next is an open-weight causal LM with a sparse MoE architecture of 80B total parameters and only ~3B activated per token, designed for coding agents and local development. It runs exclusively in non-thinking mode (no blocks emitted), ships a native 262K context window, and list OpenRouter aggregates multiple providers for this model — Parasail ($0.12/$0.80, 0.91s latency, 66 tps, 99.96% uptime), StreamLake, NovitaAI (1.00s/91 tps), Alibaba Cloud Int. (1.08s/147 tps), and Ionstream — with per-provider GPQA Diamond and TAU-Bench auto-routing benchmark splits. Notably, Cortecs is not listed amon

Videos about Qwen3 Coder Next

More models around Qwen3 Coder Next