Model details
Qwen3-Max-Thinking
Qwen3-Max-Thinking is designed for high-stakes cognitive work where deep, multi-step reasoning matters more than speed. The model leverages a scaled mixture-of-experts architecture to handle the heavy lifting of complex problem-solving, and its thinking mode allows it to strengthen reasoning steps before committing to an answer. One of its defining features is adaptive tool-use: rather than waiting for explicit user commands, it autonomously decides when to call a web search or code interpreter based on what the conversation demands, weaving those tools into its reasoning flow as needed.
The model's capabilities stem from large-scale reinforcement learning applied during training, a process that consumes substantial compute to cultivate self-correction and multi-round verification behaviors. It demonstrates particular strength on rigorous mathematical and technical reasoning tasks, achieving near-perfect scores on benchmarks like AIME 25 and HMMT when paired with test-time scaling and a code interpreter. Compared against leading models like GPT-5.2-Thinking and Claude-Opus-4.5 across nineteen benchmarks, it holds its own across factual knowledge, instruction following, alignment, and agent tasks. This makes it well-suited for technical problem-solving, research assistance, and applications where logically traceable answers are essential.
Quick Info
Powered by- Provider
- ZenMux
- Model key
- qwen/qwen3-max
- Release date
- Jan 23, 2026
- Last updated
- Jan 23, 2026
- Knowledge cutoff
- 2025-01-01
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $1.20
- Output token cost
- $6.00
Limits
- Output tokens
- 64,000 tokens
- Context window
- 256,000 tokens