Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
OpenRouter logo

Model details

Qwen3 30B A3B Instruct 2507

Qwen3 30B A3B Instruct 2507 is a mixture-of-experts Instruct refresh in the Qwen3 family, with roughly 30.5 billion total parameters but only about 3 billion activated per token, a design that aims to keep compute and latency modest while still delivering strong general assistant behavior. Unlike the earlier Qwen3-30B-A3B release, this version operates exclusively in non-thinking mode and no longer emits think blocks, which simplifies prompt design for production pipelines and downstream agents that expect direct, final-form responses. Qwen positions the model as approaching the quality of larger non-thinking variants such as Qwen3-235B-A22B, while remaining friendly to local deployment thanks to its low active-parameter footprint.

The model is intended as a balanced, instruction-tuned assistant for everyday chat, reasoning, coding, math, and multilingual workloads, with particular emphasis on alignment with open-ended user intent. Its expanded 256K context window makes it well suited for long-document summarization, codebase analysis, and multi-turn agent sessions where earlier Qwen3-30B variants would have lost track of earlier material. Weights are openly published on Hugging Face, and the same checkpoint is also served via managed inference through OpenRouter, so teams can move between self-hosted quantized builds and a hosted endpoint without changing model semantics, picking the option that best matches their cost, latency, and data-residency needs.

OpenRouterqwen/qwen3-30b-a3b-instruct-2507qwen

Quick Info

Powered by
Provider
OpenRouter
Model key
qwen/qwen3-30b-a3b-instruct-2507
Release date
Jul 29, 2025
Last updated
Jul 29, 2025
Knowledge cutoff
2025-06-30
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.04815
Output token cost
$0.19305

Limits

Output tokens
32,000 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen3 30B A3B Instruct 2507 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3 30B A3B Instruct 2507

Weights & Biases

CoverageDiscourse

This is the official HuggingFace discussions index for the canonical Qwen/Qwen3-30B-A3B-Instruct-2507 repository, confirming the exact variant exists at that path and has an active community. Open and closed threads cover topics such as recommended GPU setups for Qwen3 30B, an external OMS scoring post (OMS 70.4 B), GP Earlier threads also include user questions on static KV-cache support, Polish language coverage, LoRA fine-tuning issues, and installation guides. While the page itself is a discussion index rather than a substantive announcement, it independently corroborates the model's identity and surfaces ecosystem activity aroun

OpenRouter

Official sourceBenchmark

OpenRouter's product page for qwen/qwen3-30b-a3b-instruct-2507 confirms the model is a 30.5B-parameter mixture-of-experts language model from Alibaba's Qwen team with 3.3B active parameters per inference, operating in non-thinking mode. The page documents a 262K-token context window with a 32,000-token output cap, a kn The listing also shows current multi-provider availability through OpenRouter with five standard-latency providers — StreamLake (55% off at $0.04815 input / $0.1931 output per 1M tokens), SiliconFlow, CoreWeave, Nebius Token Factory, and Alibaba Cloud International — along with their latency, throughput, and uptime met

Weights & Biases

Coverage

This Ollama page hosts a third-party community mirror (alibayram/Qwen3-30B-A3B-Instruct-2507) of the exact Qwen variant, published about a year ago as a Q4_K_M quantization. The reproduced README describes Qwen3-30B-A3B-Instruct-2507 as a causal language model at the post-training stage with 30.5B total parameters and According to the same listing, Qwen frames this release as an updated non-thinking variant of the earlier Qwen3-30B-A3B, with stated gains in instruction following, logical reasoning, text comprehension, mathematics, science, coding, and tool usage, broader multilingual long-tail knowledge, better subjective/open-ended

Weights & Biases

Coverage

The Hugging Face model card for Qwen/Qwen3-30B-A3B-Instruct-2507 explicitly names the exact subject variant and serves as the strongest first-party technical reference among the supplied candidates. It documents the architecture as a causal language model with 30.5B total parameters and 3.3B activated parameters across The card lists enhancement highlights versus the prior Qwen3-30B-A3B release: significant gains in instruction following, logical reasoning, mathematics, science, coding, and tool usage; broader long-tail multilingual knowledge; better subjective/open-ended alignment; and enhanced 256K long-context understanding. It al

Videos about Qwen3 30B A3B Instruct 2507

More models around Qwen3 30B A3B Instruct 2507