Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Venice AI logo

Model details

Qwen 3.8 2.4T

Alibaba's flagship open-weight model pairs a 2.4 trillion parameter sparse Mixture-of-Experts architecture with a 262K token context window, positioning it as a long-horizon reasoning engine rather than a compact assistant. The sparse MoE design means only a slice of the total parameters activates per token, which lets a model this large remain practical for coding and research workflows while keeping latency reasonable. Released under a custom qwen3.8-max license rather than a standard permissive open-source license, it still qualifies as open-weight for developers who can accept those terms and want self-hostable access through Venice's private, non-surveillance infrastructure.

Independent reviewers describe the model as Alibaba's 2.4T flagship rather than a small model, with benchmark coverage highlighting strength on PaperBench and multimodal evaluations, while noting it trails Fable 5 on hard coding tasks. That mix of strengths makes it a strong fit for research synthesis, instruction following, and extended agentic sessions that benefit from the full context window, particularly when developers need sovereign control over sensitive prompts. For teams comparing options, the practical takeaway is that this model leans into long-context reasoning and coding breadth over narrow frontier coding dominance, fitting workflows where open-weight flexibility and privacy outweigh squeezing the last few percentage points on the hardest code benchmarks.

Venice AIqwen-3-8-2-4t-a95bqwen

Quick Info

Powered by
Provider
Venice AI
Model key
qwen-3-8-2-4t-a95b
Release date
Aug 12, 2026
Last updated
Aug 13, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$2.50
Output token cost
$7.50

Limits

Output tokens
65,536 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen 3.8 2.4T pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen 3.8 2.4T

Venice AI

Coverage

Alibaba released the open weights for Qwen3.8-2.4T-A95B, the open-weight variant of Qwen3.8 Max and a 2.4 trillion-parameter mixture-of-experts model with 95 billion activated parameters per token, as described in an NVIDIA Technical Blog post dated August 12, 2026. The architecture combines fine-grained MoE routing with a hybrid full and linear attention design, supports a context window of up to one million tokens, and produces outputs of up to 128K tokens for reasoning and agentic workloads. The same NVIDIA Technical Blog explains that, without additional tuning, Qwen3.8-2.4T-A95B achieves over 4,000 tokens per second per GPU and more than 350 tokens per second per user on NVIDIA GB300 NVL72 systems in FP8 precision on Day 0, with further gains expected from NVFP4 optimizations. NVIDIA NeMo AutoModel supports post-training via full supervised fine-tuning or memory-efficient LoRA, and open-source inference recipes are available for SGLang, vLLM, and NVIDIA Dynamo, with additional deployment via a model-free NVIDIA NIM container from NVIDIA NGC.

Venice AI

CoveragePreview

Alibaba's Qwen team unveiled Qwen 3.8 on July 19, 2026 via an official X post, announcing a 2.4-trillion-parameter model that the team describes as "one of the most powerful models available today" and "second only to Fable 5" — a positioning based on internal evaluations rather than third-party verification. A preview The announcement flagged Qwen 3.8 as the team's first multimodal model above the one-trillion-parameter mark, capable of processing images, videos, and documents, and expected to outperform the prior-generation Qwen 3.7-Max on coding and complex productivity tasks. Access was distributed through Token Plan subscription

Venice AI

CoverageBenchmark

Alibaba released Qwen3.8-Max on August 3, 2026, as a 2.4-trillion-parameter mixture-of-experts multimodal model with 95 billion active parameters — a configuration that corresponds to the subject model's "a95b" active-parameter designation. According to the explainer, this is Alibaba's first multimodal model above the Alibaba published a full benchmark table alongside the August 3 release, reporting PaperBench 93.0, IFBench 82.8, Terminal Bench 2.1 at 86.6, MathVision 95.2, and OSWorld-Verified 86.1, alongside weaker showings of HLE 43.6 and SWE-bench Pro 67.7 — the latter 12 points behind "Fable 5." Pricing is set at $2 per million

Videos about Qwen 3.8 2.4T

More models around Qwen 3.8 2.4T