Sulat.com
AI models
Get 10-25% off
Get 10-25% off from Qwen
Alibaba logo

Model details

Qwen3-Coder 30B-A3B Instruct

Qwen3-Coder 30B-A3B Instruct is a causal language model built on a Mixture-of-Experts architecture, featuring 128 total experts with 8 active per forward pass. Designed specifically for complex programming tasks, the model utilizes 30.5 billion parameters to deliver efficient, high-performance results. It excels in agentic coding and browser-use scenarios, supported by a specialized function-call format that integrates seamlessly with development platforms. With a native context window of 256,000 tokens—extendable to 1 million tokens using Yarn—the model is engineered for deep repository-scale understanding and structured code completion.

The model underwent comprehensive pre-training and post-training stages to refine its instruction-following capabilities. It is optimized for direct, non-thinking output, ensuring that responses remain focused and ready for immediate implementation in automated pipelines. By leveraging its expert-based design, the model maintains a balance between computational efficiency and task accuracy. Its architecture is particularly well-suited for developers requiring robust support for agentic workflows and long-context analysis, providing a reliable foundation for modern software engineering and technical automation tasks.

Alibabaqwen3-coder-30b-a3b-instructqwen

Quick Info

Powered by
Provider
Alibaba
Model key
qwen3-coder-30b-a3b-instruct
Release date
Apr 1, 2025
Last updated
Apr 1, 2025
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.45
Output token cost
$2.25

Limits

Output tokens
65,536 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen3-Coder 30B-A3B Instruct pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3-Coder 30B-A3B Instruct

Alibaba

CoverageBenchmark

Millstone AI published a third-party inference benchmark page for Qwen3-Coder-30B-A3B-Instruct in FP8 precision. The page documents the model as an FP8-quantized 30.5B-parameter Mixture-of-Experts (MoE) architecture built on the Qwen3 base, with 128 experts (8 active per forward pass) and only 3.3B parameters activated The benchmark pages report hardware-specific performance for the FP8 variant, including peak throughput of 334 tok/s on a 1x RTX Pro 6000 Blackwell (96GB), 584 tok/s on a 1x H100 SXM (80GB), and 600 tok/s on a 1x H200 SXM (141GB), tested across concurrency ranges of 1–4 to 1–6 and context lengths from 1K up to 256K. Pe

Videos about Qwen3-Coder 30B-A3B Instruct

More models around Qwen3-Coder 30B-A3B Instruct