Currently listed through these providers:
Model details
Qwen3-235B-A22B-Thinking-2507
Qwen3-235B-A22B-Thinking-2507 is a Mixture-of-Experts open-weight model built around a 235B-parameter architecture that activates just 22B parameters per forward pass. This MoE design pairs 128 expert neurons with 8 active experts per token, distributed across 94 transformer layers with grouped-query attention (64 Q heads, 4 KV heads) to keep inference efficient at scale. The model is purpose-built for thinking-only operation, with its default chat template automatically invoking a structured reasoning loop to enforce deep, step-by-step problem-solving before responding. This specialization makes it particularly effective for tasks that demand multi-step deduction, research-grade analysis, and logical consistency over conversational fluidity.
Over three months of focused development, the team scaled the model's thinking capability, refining both the depth of reasoning chains and the breadth of general skills like instruction following, tool use, and human-preference alignment. The resulting checkpoint achieves state-of-the-art performance among open-source thinking models on benchmarks spanning mathematics, science, coding, and academic reasoning that typically requires expert-level human evaluation. It also ships with notably improved long-context understanding up to 256K tokens, supporting document-level analysis and extended multi-turn problem solving. The Apache 2.0 license makes it viable for both research and production deployments, and the emphasis on reasoning quality over broad chat politeness positions it well for high-stakes applications where sound logic and verifiable chain-of-thought matter most.
Quick Info
Powered by- Provider
- Hugging Face
- Model key
- Qwen/Qwen3-235B-A22B-Thinking-2507
- Release date
- Jul 25, 2025
- Last updated
- Jul 25, 2025
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.30
- Output token cost
- $3.00
Limits
- Output tokens
- 131,072 tokens
- Context window
- 262,144 tokens
Transparent token rates
Compare Qwen3-235B-A22B-Thinking-2507 pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Qwen3-235B-A22B-Thinking-2507
No articles yet. Fetch the latest news to show it here.