Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Kilo Gateway logo

Model details

Qwen: Qwen3 30B A3B Thinking 2507 (retires Oct 9)

Qwen3-30B-A3B-Thinking-2507 is a specialized causal language model built on a Mixture-of-Experts architecture, designed to excel in tasks that demand deep logical analysis. With a total of 30.5 billion parameters, the model utilizes 128 experts and activates 8 per inference to maintain efficiency while tackling complex challenges in mathematics, coding, and scientific research. Its design centers on a dedicated thinking mode, which allows the model to generate internal reasoning traces before providing a final answer. This architecture is supported by 48 hidden layers and Group Query Attention, providing the structural foundation necessary for handling large-scale document processing and intricate multi-step reasoning tasks.

The model underwent extensive pre-training and post-training to refine its instruction following, tool usage, and alignment with human preferences. By leveraging a native context length of 262,144 tokens, it is particularly well-suited for analyzing extensive codebases and long-form documents. Its performance is marked by significant gains in academic benchmarks, such as the AIME25, where it demonstrates high-level proficiency in competitive problem solving. As an agentic-ready tool, it is engineered to integrate seamlessly into automated workflows, offering a robust solution for researchers and developers who require a balance of high-performance reasoning and efficient, scalable deployment.

Kilo Gatewayqwen/qwen3-30b-a3b-thinking-2507qwen

Quick Info

Powered by
Provider
Kilo Gateway
Model key
qwen/qwen3-30b-a3b-thinking-2507
Release date
Aug 28, 2025
Last updated
Aug 28, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.20
Output token cost
$2.40

Limits

Output tokens
32,768 tokens
Context window
81,920 tokens

Transparent token rates

Compare Qwen: Qwen3 30B A3B Thinking 2507 (retires Oct 9) pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen: Qwen3 30B A3B Thinking 2507 (retires Oct 9)

Kilo Gateway

Coverage

LM Studio's model listing for qwen3-30b-a3b-thinking-2507 describes it as an always-thinking variant of Qwen3-30B-A3B that delivers significant improvements on reasoning tasks including logical reasoning, mathematics, science, coding, and academic benchmarks. The page reports the model is a thinking-only Mixture-of-Exp According to LM Studio, the Thinking 2507 update improves long-tail knowledge coverage across multiple languages, instruction following, text generation, alignment with human preferences, and brings advanced agent capabilities with support for over 100 languages and dialects. Default sampling parameters listed are temp

Videos about Qwen: Qwen3 30B A3B Thinking 2507 (retires Oct 9)

More models around Qwen: Qwen3 30B A3B Thinking 2507 (retires Oct 9)