Currently listed through:
Model details
Qwen 3.8 Max Prime
Qwen 3.8 Max Prime is positioned as the high-speed enterprise edition within Alibaba's Qwen 3.8 generation, retaining the flagship's large-scale Mixture-of-Experts design while emphasizing faster output throughput for production environments. Source coverage describes it as built around a 2.4-trillion-parameter MoE architecture with roughly 95 billion active parameters per token, the same skeleton used across the Qwen 3.8 family, and frames Prime specifically as the variant optimized for coding, office automation, and long-running agent workflows rather than as a fundamentally separate model.
In practical terms, the model targets multimodal reasoning across text, image, and video inputs with text output, and its very long context window makes it well suited to multi-document analysis, extended agent sessions, and code repositories that exceed typical limits. It inherits the hybrid attention approach that combines Gated DeltaNet linear attention with standard Gated Attention, a design choice that supports both efficient streaming of long contexts and the precise recall needed for professional and research tasks. Together these traits make Prime a fit for teams that want frontier-class reasoning paired with higher sustained throughput in enterprise pipelines.
Quick Info
Powered by- Provider
- Kilo Gateway
- Model key
- qwen/qwen3.8-max-prime
- Release date
- Sep 23, 2026
- Last updated
- Sep 23, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $4.00
- Output token cost
- $12.00
Limits
- Output tokens
- 131,072 tokens
- Context window
- 1,000,000 tokens
Transparent token rates
Compare Qwen 3.8 Max Prime pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Qwen 3.8 Max Prime
No articles yet. Fetch the latest news to show it here.