Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Requesty logo

Model details

Qwen3.8 2.4T A95B

Qwen3.8-2.4T-A95B is a 2.4T-parameter open-weight model in the Qwen family that activates roughly 95B parameters per token, positioning it as Alibaba's largest open-weight release with near-frontier ambitions. NVIDIA chose this model to anchor serving guidance targeted at the GB300 NVL72 platform, signaling that it is intended for high-throughput, large-context agentic workloads where reasoning depth and tool integration both matter. The combination of open weights with configurable reasoning budgets lets operators tune how much deliberation the model performs per request, which is especially useful when balancing latency against answer quality in production pipelines.

Released in mid-August 2026, the model arrives alongside NVIDIA's technical documentation for running it on next-generation Blackwell infrastructure, reflecting a coordinated push to make trillion-parameter open-weight inference practical on data-center hardware. That alignment suggests strong fit for organizations building agent systems, retrieval-augmented assistants, and structured-output workflows that benefit from adjustable reasoning and long context windows. Teams selecting this model should expect a system optimized for ambitious reasoning tasks at scale rather than lightweight chat, with the operational footprint to match its parameter count and hardware recommendations.

Requestyqwen3.8-2.4T-A95Bqwen

Quick Info

Powered by
Provider
Requesty
Model key
qwen3.8-2.4T-A95B
Release date
Aug 12, 2026
Last updated
Aug 12, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$2.00
Output token cost
$6.00

Limits

Output tokens
262,144 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen3.8 2.4T A95B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.8 2.4T A95B

Requesty

CoverageBenchmark

Alibaba's Qwen team released Qwen3.8-Max, a 2.4-trillion-parameter mixture-of-experts model with 95 billion active parameters per forward pass, first via its QwenCloud API on August 2, 2026, and then as open weights packaged as Qwen3.8-2.4T-A95B on Hugging Face and ModelScope on August 13, 2026, according to this third The piece describes Qwen3.8-Max as positioned for software engineering and agent tasks, with the sparse MoE design (95B active per pass out of 2.4T total) presented as the key enabler for self-hosting without a small data center. Creator attribution to Alibaba's Qwen team is preserved, while OpenAI, Anthropic, and Goog

Requesty

Coverage

NVIDIA's technical blog documents Alibaba's open-weight release of Qwen3.8-2.4T-A95B (Qwen3.8-Max), a fine-grained mixture-of-experts model with 2.4 trillion total parameters and 95 billion activated per token, a hybrid full and linear attention architecture, a context window of up to one million tokens, and an output The post reports that on NVIDIA GB300 NVL72 hardware in FP8 precision, Qwen3.8-2.4T-A95B achieves over 4,000 tokens per second per GPU and over 350 tokens per second per user on Day 0, with further gains expected from NVFP4 optimizations. NVIDIA NeMo AutoModel supports post-training via full supervised fine-tuning or m

Videos about Qwen3.8 2.4T A95B

More models around Qwen3.8 2.4T A95B