Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Opper logo

Model details

Qwen3.8 2.4T A95B

Qwen3.8 2.4T A95B is positioned as the flagship release of Qwen's 3.8 generation and represents the largest model the team has shipped to date. According to a Together AI model page, it is built as a sparse Mixture-of-Experts design with 2.4 trillion total parameters, and Qwen describes it as their first multimodal model above one trillion parameters, aimed at coding, long-horizon agentic work, and multimodal reasoning. The launch was associated with an open-weight commitment, with Hugging Face repository links surfaced in a Hacker News discussion thread referencing bf16 and FP8 releases available at launch.

The intended use profile combines very long contexts with always-on reasoning for complex, multi-step tasks. Together AI describes a million-token context window with output capacity into the hundreds of thousands of tokens, and notes thinking is permanently enabled with adjustable reasoning effort across low, high, and xhigh levels, letting callers trade response speed against deliberation depth. Practical fit centers on repository-scale code analysis, long-horizon agentic pipelines, and multimodal reasoning scenarios where sustained chain-of-thought and substantial output budgets matter more than lightweight chat use.

Opperqwen3.8-2.4t-a95bqwen

Quick Info

Powered by
Provider
Opper
Model key
qwen3.8-2.4t-a95b
Release date
Aug 12, 2026
Last updated
Aug 12, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$2.50
Output token cost
$6.00

Limits

Output tokens
131,072 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen3.8 2.4T A95B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.8 2.4T A95B

Opper

CoverageBenchmark

Shattered.io reports that Qwen3.8-Max, a 2.4-trillion-parameter mixture-of-experts model with 95 billion active parameters, first went live through Alibaba's QwenCloud API on August 2, 2026, with the open-weight variant Qwen3.8-2.4T-A95B following on Hugging Face and ModelScope on August 13 under an Apache 2.0 license. The article positions the open-weight Qwen3.8 release within a crowded September for frontier AI, presenting benchmark context including an 86.6 score claimed to position it near Opus 5, though the supplied excerpt does not cite a primary benchmark source for that figure. Attribution to a "Laura Bennett" byline on a th

Opper

Coverage

Alibaba released the open weights for Qwen3.8-2.4T-A95B, its largest open-weight model, as a 2.4 trillion total parameter mixture-of-experts architecture with 95 billion activated parameters per token, a hybrid full and linear attention design, a context window of up to one million tokens, and an output length of up to The article confirms the model is released under an Apache 2.0 license and is positioned as a near-frontier open alternative to closed APIs, aligning the open-weight variant with Qwen3.8-Max in the model card's framing. For developers, it provides actionable deployment paths: downloading weights from Hugging Face for l

Opper

CoverageDiscourse

A community discussion on the Qwen/Qwen3.8-2.4T-A95B Hugging Face repository highlights that the released open weights are text-only and lack several capabilities of the hosted Qwen3.8-Max variant, specifically vision input and a native 1M context window. The thread quotes the official model card, which states that Qwe Beyond the feature gap, the discussion expresses frustration about vision-language capability being withheld while 1M context has become an industry standard, and draws unfavorable comparisons to Kimi K3's full release and an anticipated Muse Spark 1.2 open-weight release. The post is user opinion rather than an offici

Videos about Qwen3.8 2.4T A95B

More models around Qwen3.8 2.4T A95B