Currently listed through these providers:
Model details
Qwen3 Max
Qwen3-Max is the largest model in the Qwen3 family, built on the same MoE Mixture-of-Experts architecture used across the series and incorporating a global-batch load balancing loss to keep pretraining stable. It was pretrained on 36 trillion tokens and exceeds 1 trillion parameters, giving it substantial capacity for knowledge-intensive and reasoning-heavy workloads. The scaling approach continues the Qwen3 design paradigm, and a Thinking variant is still under active training and not yet publicly released.
Intended for demanding applications such as long-context reasoning, coding, instruction following, and multilingual tasks, Qwen3-Max-Instruct is positioned as a production-ready text model. According to Alibaba, the preview ranked third on the Text Arena leaderboard, and the official release is reported to achieve state-of-the-art results across benchmarks covering knowledge, reasoning, coding, instruction following, human preference alignment, agent tasks, and multilingual understanding. The companion Thinking variant, when augmented with tool usage and scaled test-time compute, has reportedly reached 100% on AIME 25 and HMMT, suggesting strong potential for advanced agent and reasoning workflows once it becomes broadly available.
Quick Info
Powered by- Provider
- Deep Infra
- Model key
- Qwen/Qwen3-Max
- Release date
- Sep 23, 2025
- Last updated
- Sep 23, 2025
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $1.20
- Output token cost
- $6.00
Limits
- Output tokens
- 65,536 tokens
- Context window
- 256,000 tokens