Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Alibaba Token Plan (China) logo

Model details

Qwen3.8 Max

Alibaba's Qwen3.8 Max is a sparse mixture-of-experts model with roughly 2.4 trillion total parameters and about 95 billion active per forward pass, an architecture that keeps inference tractable while supporting a very large overall capacity. It was first exposed through Alibaba's QwenCloud API, then published as the open-weight Qwen3.8-2.4T-A95B checkpoint on Hugging Face and ModelScope under an Apache 2.0 license, giving well-resourced teams the option to self-host rather than rent access. A 27-billion-parameter dense sibling shipped open weights a day later for teams that cannot provision hardware for the full MoE, while the hosted Qwen3.8-Max endpoint remains the route for reaching the larger configuration directly.

Qwen3.8 Max is positioned for software engineering and long-horizon agent work rather than general chat. On Terminal-Bench 2.1 it lands within roughly two points of top closed frontier systems, and on the updated Qwen3.8-Max-0902 build it closes additional ground on agent-focused suites such as DeepSWE and MLS-Bench-Lite. Independent benchmark tracking gives the model a top-five placement for multimodal and grounded tasks and a top-ten placement for agentic tool use, with notably strong performance on screenshots, charts, and document comprehension. That balance of coding, reasoning, and multimodal grounding makes it a natural fit for coding agents and retrieval-augmented workflows where a self-hostable alternative to closed frontier APIs is desired.

Alibaba Token Plan (China)qwen3.8-maxqwen

Quick Info

Powered by
Provider
Alibaba Token Plan (China)
Model key
qwen3.8-max
Release date
Aug 3, 2026
Last updated
Aug 3, 2026
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
131,072 tokens
Context window
1,000,000 tokens

Latest news about Qwen3.8 Max

Alibaba Token Plan (China)

CoverageBenchmark

Qwen3.8-Max is described in this write-up as a 2.4-trillion-parameter mixture-of-experts model with 95 billion active parameters per forward pass, positioning it for coding and multi-step agent workloads rather than general chat. The piece states that Alibaba first made the model available through its QwenCloud API on The article frames the rollout as part of a broader August–September 2026 push by Alibaba alongside other frontier releases, and highlights the Apache 2.0 licensing choice for the open-weight variant as a signal in the open-versus-closed model debate. It cites a headline benchmark figure of 86.6 in connection with Qwen

Alibaba Token Plan (China)

CoverageBenchmark

This BenchLM.ai aggregator page tracks Qwen3.8 Max across multiple benchmark categories, reporting an aggregate capability score of 71.8/100 against a field median of 56.1 and ranking it 10th of 233 tracked models. It lists 1M-token context as the maximum window (separate from a tracked maximum output length) and a fir The page breaks down verified evidence by category, with 15/15 agentic, 12/12 coding, 2/2 reasoning, and 6/6 knowledge benchmarks marked verified, while math and multilingual categories are listed as not measured; it reports a reasoning score of 86.6 and a knowledge score of 68.8, with rankings in the 88th–95th percent

Videos about Qwen3.8 Max

More models around Qwen3.8 Max