Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Fireworks AI logo

Model details

Qwen3.8 2.4T A95B

Qwen3.8 2.4T A95B is the open-weight counterpart to Alibaba's Qwen3.8-Max flagship, structured as a sparse mixture-of-experts language model with roughly 2.4 trillion total parameters and about 95 billion activated per token. The design uses fine-grained routing across 512 experts and pairs full-attention layers with linear-attention layers, a combination intended to keep compute and key-value cache growth in check as inputs stretch toward very long contexts. Weights were published on Hugging Face under the Qwen organization, giving researchers and operators direct access to the same architecture that drives the hosted Max-tier model rather than a distilled or trimmed variant.

In practical terms the model is aimed at coding, research, and long-horizon agentic workflows where sustained reasoning and tool use matter more than lightweight chat. The sparse routing concentrates capacity into the experts most relevant to each token, which is well suited to multi-step problem solving and large code or document analyses. Because it is a text-only model and its full parameter count requires data-center class hardware, the natural fit is for teams that need frontier-scale reasoning on their own infrastructure or through providers that can host the dense expert footprint, while smaller experimental deployments are limited to aggressive quantizations that trade quality for footprint.

Fireworks AIaccounts/fireworks/models/qwen3p8-2p4t-a95bqwen

Quick Info

Powered by
Provider
Fireworks AI
Model key
accounts/fireworks/models/qwen3p8-2p4t-a95b
Release date
Aug 12, 2026
Last updated
Aug 12, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$2.00
Output token cost
$6.00

Limits

Output tokens
131,072 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen3.8 2.4T A95B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.8 2.4T A95B

Fireworks AI

Coverage

A Japanese technical explainer (AI-translated to English) from AI-Driven Lab covers the Qwen3.8 series and explicitly names the 2.4T-A95B variant alongside Qwen3.8-Max branding, describing it as Alibaba's largest model with 2.4 trillion total parameters and 9.5 billion active parameters announced on August 3, 2026. The The piece adds licensing nuance, noting that while Qwen3.8-Max weights are released as open weight, the license is "open weight but not unconditionally free," a practical consideration for developers evaluating adoption. Because the translation is AI-generated and the source is a secondary commentary rather than an Ali

Fireworks AI

Coverage

NVIDIA published an official technical blog on August 12, 2026 announcing Day-0 infrastructure support for Alibaba's Qwen3.8-2.4T-A95B open-weight model on GB300 NVL72 systems. The post confirms the model as a 2.4 trillion parameter fine-grained mixture-of-experts with 95 billion activated parameters per token, combini The blog reports measured throughput exceeding 4,000 tokens per second per GPU and over 350 tokens per second per user on NVIDIA GB300 NVL72 in FP8 precision, with further gains expected from upcoming NVFP4 optimizations. NVIDIA NeMo AutoModel supports post-training via full supervised fine-tuning or memory-efficient L

Fireworks AI

CoverageDiscourse

A community discussion opened on the Qwen/Qwen3.8-2.4T-A95B Hugging Face repository documents significant gaps between the open-weight release and the hosted Qwen3.8-Max service. According to the thread, the released weights are text-only and do not include the vision input capability that is available in the hosted ve The discussion further notes that the open weights omit the official built-in tools present in the hosted Qwen3.8-Max, framing these omissions as a developer-experience concern for users who expected feature parity with the hosted model. The thread is sourced from the model's own Hugging Face repository, making it dire

Fireworks AI

CoverageBenchmark

Artificial Analysis provides an independent benchmark and technical reference for the exact Qwen3.8 2.4T A95B variant, reporting a score of 47 on the Artificial Analysis Intelligence Index (placing it above the comparable-model median of 22), 2.4T total parameters with 95B active per token, text input/output modalities Independent measurements show 40.3 output tokens per second, classifying the variant as notably slow relative to peers, and report total evaluation cost of $1,779.87 on the Intelligence Index alongside 150M output tokens generated. These third-party benchmark findings give developers a quantitative signal about the mod

Videos about Qwen3.8 2.4T A95B

More models around Qwen3.8 2.4T A95B