Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
TensorX logo

Model details

Qwen3.8 2.4T A95B

Qwen3.8 2.4T A95B is an open-weight mixture-of-experts model released by Alibaba with 2.4 trillion total parameters and 95 billion activated per token. Its architecture pairs full attention layers with linear-attention layers, a fine-grained expert design that keeps compute and memory bounded as context scales, and a window reaching up to one million tokens. Built-in configurable reasoning controls let developers tune inference depth per request, making the model well suited to coding, large-scale document analysis, and long-running agentic workflows where context grows with tool outputs, retrieved passages, and multi-step traces.

Out of the box on NVIDIA GB300 NVL72 in FP8 precision, the model delivers over 4,000 tokens per second per GPU and over 350 tokens per second per user, with NVFP4 expected to push throughput further. The ecosystem around it is broad: Hugging Face hosts the official weights, NVIDIA NeMo AutoModel enables post-training via full supervised fine-tuning or memory-efficient LoRA on those checkpoints, and open-source inference recipes ship for SGLang, vLLM, and NVIDIA Dynamo, with a model-free NIM container available from NGC. The combination of a trillion-parameter sparse design, hybrid attention for long context, and a mature serving and fine-tuning stack positions the model as a flexible foundation for research and production agentic systems.

TensorXqwen/qwen3.8-2.4t-a95bqwen

Quick Info

Powered by
Provider
TensorX
Model key
qwen/qwen3.8-2.4t-a95b
Release date
Aug 12, 2026
Last updated
Aug 12, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$2.50
Output token cost
$6.00

Limits

Output tokens
64,000 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen3.8 2.4T A95B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.8 2.4T A95B

TensorX

CoverageBenchmark

Shattered.io provides a release timeline for the Qwen3.8 family, reporting that Qwen3.8-Max first became available through Alibaba's QwenCloud API on August 2, 2026, and that the open-weight variant Qwen3.8-2.4T-A95B followed on Hugging Face and ModelScope on August 13, 2026, with a dense sibling, Qwen3.8-27B, releasin The piece situates Qwen3.8-2.4T-A95B within the broader competitive landscape, citing secondary sources such as Latent.Space (August 3) and CheapestInference for release details and framing the choice between renting a frontier API and running one's own weights. While the headline score of 86.6 is reported, the article

TensorX

CoverageBenchmark

Lyceum Technology describes Qwen3.8 2.4T A95B as an open-weight Mixture-of-Experts flagship with 2.4 trillion total parameters and 95 billion active parameters per token, built around a fine-grained MoE topology that distributes capacity across 512 routed experts and a shared expert module, activating 10 routed experts Lyceum frames Qwen3.8 2.4T A95B as an "open-weight foundation for complex agentic workflows," arguing the architecture keeps per-step compute close to sub-100B-class footprints while still delivering frontier-level reasoning depth for multi-step code execution, autonomous tool orchestration, and long-context reasoning.

TensorX

Coverage

NVIDIA's technical blog confirms that Alibaba released the open weights for Qwen3.8-2.4T-A95B, a sparse mixture-of-experts model with 2.4 trillion total parameters and 95 billion activated per token. The model uses a fine-grained MoE architecture with a hybrid of full and linear attention, supports a context window of The blog documents the post-training and serving tooling available for the exact Qwen3.8-2.4T-A95B variant: NVIDIA NeMo AutoModel supports full supervised fine-tuning and memory-efficient LoRA fine-tuning directly on Hugging Face checkpoints for domain-specific use cases, and open-source inference recipes are available

TensorX

CoverageBenchmark

Artificial Analysis publishes a benchmark and price analysis for Qwen3.8 2.4T A95B, identifying it as an open-weight mixture-of-experts model from Alibaba released in August 2026 with 2.4 trillion total parameters, 95 billion active per token, a 984k-token context window, text in/out, and a reasoning variant. The page The analysis positions Qwen3.8 2.4T A95B within Artificial Analysis's 113-model evaluation set, ranking it 4th on Intelligence, 51st on Speed, and 35th on Cost, and compares it against same-class open-weight peers. Pricing is flagged as expensive relative to a $0.30 input and $1.15 output median, and evaluation of the

Hugging Face

CoverageBenchmark

OpenRouter’s listing identifies Qwen3.8 2.4T A95B as a Qwen open-weight sparse mixture-of-experts model with 2.4 trillion total parameters and 95 billion active parameters. It lists a 1-million-token context window and describes intended use across coding, research, complex reasoning, and agentic workflows. The same page records an August 12, 2026 release date and says the model accepts and produces text. The remainder of the excerpt is a gateway dashboard focused on provider routing, pricing, latency, throughput, uptime, and provider-specific benchmark results, so those hosting details should not be treated as model-leve

Videos about Qwen3.8 2.4T A95B

More models around Qwen3.8 2.4T A95B