LLM Gateway
Alibaba debuts Qwen3-Max, its trillion-parameter AI model trained on 36T tokens. The system handles 1M-token inputs and is available through Alibaba Cloud.
Model details
Qwen3-Max represents the most advanced iteration in Alibaba's Qwen3 series, built on a Mixture of Experts architecture with over one trillion parameters. The model was pretrained on 36 trillion tokens and incorporates a global-batch load balancing loss function during training. This design enables the system to process extremely long contexts up to one million tokens while maintaining stable training dynamics and smooth loss curves. The architecture delivers state-of-the-art performance across knowledge, reasoning, coding, instruction following, and multilingual understanding.
The release version builds on the preview's strong foundation, ranking third on the Text Arena leaderboard where it surpassed GPT-5-Chat. The Thinking variant leverages substantial reinforcement learning compute to push performance even further, achieving 100% accuracy on challenging reasoning benchmarks like AIME 25 and HMMT through advanced test-time scaling techniques. Adaptive tool-use capabilities allow on-demand retrieval and code interpreter invocation, making it particularly suited for complex agent workflows. Benchmarks show it competing with leading models like GPT-5.2-Thinking, Claude-Opus-4.5, and Gemini 3 Pro across 19 established tests, positioning it as a strong choice for applications requiring both deep reasoning and autonomous task execution.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
LLM Gateway
Alibaba debuts Qwen3-Max, its trillion-parameter AI model trained on 36T tokens. The system handles 1M-token inputs and is available through Alibaba Cloud.
LLM Gateway
🎯 Key Takeaways (TL;DR) Breakthrough Scale: Alibaba releases first trillion-parameter... Tagged with qwen3.