Currently listed through these providers:
Model details
Qwen/Qwen3-30B-A3B-Instruct-2507
Qwen3-30B-A3B-Instruct-2507 is a Mixture of Experts language model built around a sparse activation pattern that keeps inference efficient while scaling up parameter count. With 128 distinct experts in the network and only 8 of them engaged per forward pass, the model directs specialized capacity toward different types of problems without activating the full 30.5 billion parameters simultaneously. This architecture rests on 48 transformer layers using group query attention, where a small set of key-value heads serves many query heads, reducing memory demands during long-context work. The design intent is clear: provide strong capability in instruction following, coding, mathematics, and reasoning tasks, while keeping the compute footprint practical enough for real-world deployment at scale.
The model was developed through a pretraining and post-training pipeline, with the July 2025 update (designated "2507") marking significant gains across the board. Post-training refinements sharpened its instruction adherence, logical reasoning, and tool-use competency, while substantially improving how it aligns with user preferences on open-ended and subjective tasks. Multilingual coverage received particular attention, with better retention of long-tail knowledge across languages. A defining characteristic of this specific variant is its exclusive non-thinking operation: it produces direct responses without generating intermediate reasoning blocks, making it well suited for applications where response speed and clarity matter more than step-by-step explanation. The combination of long-context understanding up to 256K tokens and efficient sparse activation positions it as a practical choice for complex document processing, code generation, and multi-turn assistant workloads where the full MoE capacity can be leveraged selectively.
Quick Info
Powered by- Provider
- SiliconFlow (China)
- Model key
- Qwen/Qwen3-30B-A3B-Instruct-2507
- Release date
- Jul 30, 2025
- Last updated
- Nov 25, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.09
- Output token cost
- $0.30
Limits
- Output tokens
- 262,000 tokens
- Context window
- 262,000 tokens
Transparent token rates
Compare Qwen/Qwen3-30B-A3B-Instruct-2507 pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Qwen/Qwen3-30B-A3B-Instruct-2507
No articles yet. Fetch the latest news to show it here.