Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Vercel AI Gateway logo

Model details

Qwen3.8 2.4T A95B

Alibaba released Qwen3.8-2.4T-A95B on August 12, 2026 as the open-weight counterpart to its Qwen3.8-Max flagship, positioning it as the largest open release from the Qwen team to date. The architecture is a fine-grained sparse mixture-of-experts design, carrying 2.4 trillion total parameters while activating roughly 95 billion per token, with computation routed across 512 experts. To keep compute and key-value cache memory manageable as inputs lengthen, the model interleaves standard attention layers with linear-attention layers, which is one of the more distinctive structural choices in this scale class.

The intended workload mix leans toward coding, research assistance, and long-running agentic tasks, with the mixture-of-experts routing and long-context behavior aimed at sustained multi-step workflows rather than single-turn chat. Community discussion noted the model is a conceptual rival to Kimi K3 in scale, though at launch only bf16 and fp8 builds were available, making data-center hardware a practical requirement, while quantized community builds such as a roughly 397 GB 1-bit dynamic variant enable more constrained local experimentation. The available weights are text-only, so vision and the full million-token context tied to the hosted variant are not part of the downloadable artifact.

Vercel AI Gatewayalibaba/qwen3.8-2.4t-a95bqwen

Quick Info

Powered by
Provider
Vercel AI Gateway
Model key
alibaba/qwen3.8-2.4t-a95b
Release date
Aug 12, 2026
Last updated
Aug 12, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$2.00
Output token cost
$6.00

Limits

Output tokens
128,000 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen3.8 2.4T A95B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.8 2.4T A95B

Vercel AI Gateway

CoverageBenchmark

Qwen3.8-Max, a 2.4-trillion-parameter mixture-of-experts model with 95 billion active parameters, first went live through Alibaba's QwenCloud API on August 2, 2026, with the open-weight variant Qwen3.8-2.4T-A95B following on August 13 on Hugging Face and ModelScope, and a smaller dense sibling Qwen3.8-27B shipping open The article situates the Qwen3.8-2.4T-A95B open-weight release within a crowded September for frontier AI, noting OpenAI, Anthropic, and Google all shipped new models within days of each other. The sparse MoE design, with only 95 billion parameters active on any forward pass, is highlighted as what makes self-hosting f

Vercel AI Gateway

Coverage

"Qwen3.8-Max," announced by Alibaba on August 3, 2026, is the company's largest model to date, with 2.4 trillion total parameters. The AI-Driven Lab guide on note.com explains the relationship between the three Qwen3.8 variants — Max, 2.4T-A95B, and 27B — detailing the underlying architecture that combines MoE with hyb The guide emphasizes license nuance, characterizing the release as "open weight but not unconditionally free," and is aimed at engineers and product managers evaluating the Qwen3.8 series for implementation. As an auto-translated explainer it carries an explicit accuracy/nuance caveat, and an active-parameter figure of

Vercel AI Gateway

CoverageBenchmark

WhatLLM's model page for Qwen3.8 2.4T A95B confirms the architecture as 2.4T total / 95B active parameters, a native 262,144-token context extensible to 1,010,000 tokens, mandatory thinking mode with low, medium, and xhigh effort levels, and weights published on the official Qwen Hugging Face organization. Independent Observed deployment metrics across hosted endpoints include fastest observed throughput of 125.1 tokens per second and a lowest observed time-to-first-token of 0.648s, with listed pricing around $1.90 per million input and $5.70 per million output tokens. The analysis positions A95B as an open-weight governance and cus

Vercel AI Gateway

Coverage

Alibaba released the open weights for Qwen3.8-2.4T-A95B (also called Qwen3.8-Max), its largest open-weight model, featuring 2.4 trillion total parameters with 95 billion activated per token in a fine-grained mixture-of-experts architecture with hybrid full and linear attention, a context window of up to one million tok The post details the deployment ecosystem around Qwen3.8-2.4T-A95B: NVIDIA NeMo AutoModel supports post-training with full supervised fine-tuning or memory-efficient LoRA fine-tuning directly on Hugging Face checkpoints, while open-source inference recipes are available for SGLang, vLLM, and NVIDIA Dynamo. The model ca

Vercel AI Gateway

CoverageDiscourse

A Hugging Face community discussion on the Qwen/Qwen3.8-2.4T-A95B model repository highlights that the released open weights are text-only and lack several features available in the hosted Qwen3.8-Max variant, notably vision input and the native 1-million-token context window. The thread quotes the official model card The discussion raises concerns about feature gating between the open-weight release and the managed Qwen Cloud deployment, noting that vision-language capabilities and extended context are reserved for the closed Max version. While the open Qwen3.8-2.4T-A95B weights remain useful for general tasks and distillation, the

Vercel AI Gateway

CoverageBenchmark

Alibaba's Qwen team open-sourced the Qwen3.8-2.4T-A95B model, marking the first time the company has publicly released weights for a Max-level flagship model. The model has 2.4 trillion total parameters with 95 billion activated per token, natively supports 262,144 tokens of context extendable to 1.01 million tokens, a The model adopts the underlying architecture of Qwen3.5 and achieves performance improvements in programming, office work, scientific research, and long-cycle agent tasks, deployable through SGLang, vLLM, and TokenSpeed inference engines with framework-specific parallel strategies based on weight precision and GPU conf

Videos about Qwen3.8 2.4T A95B

More models around Qwen3.8 2.4T A95B