Kilo Gateway
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Model details
Qwen3-Next 80B-A3B Instruct is described by third-party listings as an 80-billion-parameter sparse mixture-of-experts language model with roughly 3 billion parameters active per token, paired with a hybrid attention design that combines linear Gated DeltaNet layers with standard gated attention. That combination is presented as a deliberate trade between long-context efficiency and standard transformer recall, making the model well suited to extended-document reasoning, multilingual prompts, and coding sessions that benefit from deep context retention. The instruction-tuned variant focuses on practical assistant behavior, so it is positioned less as a base model and more as a ready-to-use conversational and task-following system for developers who want strong long-context handling without retraining.
Independent leaderboard tracking places Qwen3-Next 80B-A3B Instruct in the middle of the pack overall, with average-tier placements in legal, finance, chat, and healthcare categories, while ranking below the top half of current open models on reasoning, coding, tool calling, vision, writing, and math tasks. In practice this means it reads as a balanced generalist rather than a specialist: a sensible default when long context, open weights, and multilingual coverage matter more than top-tier benchmark scores, and a less compelling pick for workloads that demand the strongest reasoning or coding accuracy. Its sparse MoE design also suggests a favorable cost-to-quality profile at inference time, aligning with its role as an accessible long-context assistant for everyday production use.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Kilo Gateway
We’re on a journey to advance and democratize artificial intelligence through open source and open science.