Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
LMStudio logo

Model details

Qwen3 Coder 30B

Qwen3 Coder 30B is a coding-focused language model from Alibaba's Qwen family, designed around an agentic programming workflow. It uses a Mixture-of-Experts architecture with 30 billion total parameters but only 3 billion active per token, giving it the footprint of a small model while drawing on the knowledge of a much larger one. The model's full name is Qwen3-Coder-30B-A3B-Instruct, and it is one of two sizes in the Qwen3-Coder line, sitting beside a larger 480B/35B-active variant for teams that need more capacity.

For practical use, Qwen3 Coder 30B is built to handle repository-scale work, with native support for a 256K token context window that lets it reason across large files and multiple modules at once. It supports tool use and ships in both GGUF and MLX formats, so it can run locally on a wide range of consumer hardware; the LM Studio package is roughly 15 GB and needs about 15 GB of RAM to load. The model fits developers who want a self-hosted coding assistant for agentic tasks such as multi-step editing, browser-based automation, and tool-driven workflows, without depending on a remote API.

LMStudioqwen/qwen3-coder-30bqwen

Quick Info

Powered by
Provider
LMStudio
Model key
qwen/qwen3-coder-30b
Release date
Jul 23, 2025
Last updated
Jul 23, 2025
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
65,536 tokens
Context window
262,144 tokens

Latest news about Qwen3 Coder 30B

LMStudio

CoverageBenchmark

Millstone AI published a detailed inference benchmark for Qwen3-Coder-30B-A3B-Instruct in FP8 precision, documenting the model's architecture and hardware performance. The page describes it as a 30.5B-parameter Mixture-of-Experts model with 128 experts (8 active per forward pass) and 3.3B activated parameters at runtim The benchmark reports peak throughput of 334 tok/s on a single RTX Pro 6000 Blackwell 96GB, 584 tok/s on a single H100 SXM 80GB, and 600 tok/s on a single H200 SXM 141GB, with tested concurrency up to 6 and context lengths ranging from 1K up to 256K. Capacity-planning tables indicate practical concurrent request counts

LMStudio

CoverageDiscourse

A Hugging Face community discussion opened October 22, 2025, asks the Qwen team for reproducibility details behind the reported 51.6 score for Qwen3-Coder-30B-A3B-Instruct on SWE-bench Verified using OpenHands scaffolding. The user specifically requests the OpenHands commit/version, full config.toml and CLI command, th The thread remains a question without an official answer captured in the excerpt, highlighting ongoing community uncertainty about the exact benchmark configuration. It signals that the 51.6 SWE-bench Verified figure, while widely cited, lacks publicly documented reproduction specifics for the 30B-A3B Instruct variant.

Videos about Qwen3 Coder 30B

More models around Qwen3 Coder 30B