Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
OpenRouter logo

Model details

Qwen3 30B A3B Thinking 2507

Qwen3-30B-A3B-Thinking-2507 is a Mixture-of-Experts reasoning model built on a causal language model architecture with 128 total experts, 48 transformer layers, and group-based query attention. Only 3.3 billion of its 30.5 billion total parameters activate during inference, enabling efficient computation while maintaining strong reasoning capability. The model is purpose-built for thinking mode, where internal reasoning traces are separated from final outputs, and it achieves this without requiring an explicit enable_thinking flag. Its extended thinking length makes it particularly suited for tackling highly complex problems that demand deep, multi-step reasoning.

The model's training lineage traces to a deliberate scaling of Qwen3's thinking capability over several months, improving both the quality and depth of reasoning on tasks like mathematics, science, coding, and academic benchmarks that typically require human expertise. Beyond specialized reasoning, it has enhanced general capabilities including instruction following, tool usage, text generation, and alignment with human preferences. The architecture supports native understanding of very long contexts, enabling it to maintain coherence across extended inputs. Available as open weights with tool calling support and only 17GB minimum system memory, it offers a capable reasoning option for developers who want to run reasoning-intensive workloads locally or through API access.

OpenRouterqwen/qwen3-30b-a3b-thinking-2507qwen

Quick Info

Powered by
Provider
OpenRouter
Model key
qwen/qwen3-30b-a3b-thinking-2507
Release date
Aug 28, 2025
Last updated
Aug 28, 2025
Knowledge cutoff
2025-06-30
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.20
Output token cost
$2.40

Limits

Output tokens
32,768 tokens
Context window
81,920 tokens

Transparent token rates

Compare Qwen3 30B A3B Thinking 2507 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3 30B A3B Thinking 2507

OpenRouter

Coverage

The LM Studio model page for Qwen3-30B-A3B-Thinking-2507 documents the model's exact variant with concrete technical specifications. It is an always-thinking mode Mixture-of-Experts model featuring significant improvements in reasoning tasks including logical reasoning, mathematics, science, coding, and academic benchm The page further notes substantial gains in long-tail multilingual knowledge coverage, improved alignment with user preferences, and advanced agent capabilities supporting over 100 languages and dialects. It is distributed in GGUF and MLX (4-bit, 6-bit, 8-bit) formats via lmstudio-community on Hugging Face, with a mini

Videos about Qwen3 30B A3B Thinking 2507

More models around Qwen3 30B A3B Thinking 2507