Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
iFlow logo

Model details

Qwen3-Max

Qwen3-Max represents Alibaba's flagship large language model built on a Mixture of Experts architecture that distributes computation across specialized pathways, enabling efficient handling of over a trillion total parameters. The model's design incorporates a global-batch load balancing loss that stabilizes training dynamics, resulting in consistently smooth pretraining loss curves throughout the process. This architectural foundation supports state-of-the-art performance across knowledge tasks, reasoning challenges, instruction following, human preference alignment, and multilingual understanding. The model handles million-token inputs and was specifically engineered to compete with leading frontier systems in capability and benchmark standing.

The model underwent pretraining on 36 trillion tokens, an exceptionally large corpus that underpins its broad competency range. Its preview version achieved third place on the Text Arena leaderboard, surpassing GPT-5-Chat, while its coding capabilities reached a SWE-Bench Verified score of 69.6. A dedicated Thinking variant extends the base model with test-time compute scaling that iteratively refines reasoning by drawing on prior inference steps, enabling it to autonomously invoke tools like Search, Memory, and Code Interpreter during conversation. This reasoning model achieved perfect scores on mathematical benchmarks and outperformed competing frontier systems on the Humanity's Last Exam by a significant margin. The combination of massive pretraining scale, adaptive reasoning modes, and tool autonomy positions Qwen3-Max for complex agentic workflows and demanding professional applications.

iFlowqwen3-maxqwen

Quick Info

Powered by
Provider
iFlow
Model key
qwen3-max
Release date
Jan 1, 2025
Last updated
Jan 1, 2025
Knowledge cutoff
2024-12
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
32,000 tokens
Context window
256,000 tokens

Latest news about Qwen3-Max

iFlow

Coverage

Alibaba has released Qwen3-Max, a large language model with over 1 trillion parameters, which aims to compete with leading AI models like GPT-5.

iFlow

CoverageBenchmark

OpenRouter lists Qwen3 Max as an updated release on the Qwen3 series, dated September 23, 2025 on the aggregator, with a knowledge cutoff of June 2025 and a 262K-token context window. The listing advertises major improvements over the January 2025 Qwen3-Max in reasoning, instruction following, multilingual coverage (ov On OpenRouter, the model is hosted exclusively by Alibaba Cloud International — every request is forwarded directly to that single provider — and reported performance figures include a P50 latency of 0.98–0.99 seconds and throughput of roughly 38–48 tokens per second (P50), with a 43.8% cache hit rate. Cache pricing ti

Videos about Qwen3-Max

More models around Qwen3-Max