Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Kilo Gateway logo

Model details

Qwen3-Coder 30B-A3B Instruct

Qwen3-Coder-30B-A3B-Instruct is an open-source language model created by Qwen and distributed under the Apache 2.0 license. It is built on the Qwen3 architecture and uses a Mixture-of-Experts design with 30.5B total parameters, 128 experts, and 8 active experts per forward pass, activating only about 3.3B parameters during inference. The transformer has 48 layers and uses grouped-query attention with 32 query heads and 4 key-value heads, giving it a strong balance between capacity and compute efficiency for its size class.

The model is designed primarily for agentic coding work, including repository-scale code understanding, browser-based interactions, and structured tool calling through OpenAI-compatible formats. It natively supports a 256K token context window, which can be extended up to 1M tokens with YaRN, making it well suited for fitting large codebases or long technical documents into a single prompt. A fine-grained FP8 quantized variant is also available, operating in non-thinking mode for efficient deployment, and the model ships with sharded safetensor weights that fit on high-memory single-node GPU setups.

Kilo Gatewayqwen/qwen3-coder-30b-a3b-instructqwen

Quick Info

Powered by
Provider
Kilo Gateway
Model key
qwen/qwen3-coder-30b-a3b-instruct
Release date
Apr 1, 2025
Last updated
Apr 1, 2025
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.2925
Output token cost
$1.4625

Limits

Output tokens
235,929 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen3-Coder 30B-A3B Instruct pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3-Coder 30B-A3B Instruct

OpenRouter

Coverage

LLM Explorer lists Qwen3 Coder 30B A3B Instruct as an open-source model maintained by Qwen, with a reported VRAM footprint of 61.1 GB at full precision and a 256K context length under the Apache-2.0 license. The page records the Hugging Face repository as the source of truth, names the architecture as Qwen3MoeForCausal The model is sharded across sixteen roughly 4.0 GB safetensors files plus a final 1.1 GB shard, stored in bfloat16, which is consistent with self-hosted deployment of the full-precision variant. LLM Explorer also catalogs community quantized derivatives of Qwen3 Coder 30B A3B Instruct, including AWQ 4-bit, GPTQ 4-bit,

OpenRouter

CoverageBenchmark

Millstone AI published an inference benchmark page for the FP8-quantized variant of Qwen3-Coder-30B-A3B-Instruct, confirming it is a 30.5B parameter MoE model with 128 experts and 8 active per forward pass, activating only 3.3B parameters at inference time. The page documents 48 layers with GQA (32 query heads, 4 KV he The benchmark reports measured throughput and capacity across hardware: 334 tok/s on a single RTX Pro 6000 Blackwell 96GB, 584 tok/s on a single H100 SXM 80GB, and 600 tok/s on a single H200 SXM 141GB, with concurrency and context-length ranges detailed per device. Capacity planning tables list concurrent request limit

Videos about Qwen3-Coder 30B-A3B Instruct

More models around Qwen3-Coder 30B-A3B Instruct