Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Cortecs logo

Model details

Qwen3.8 2.4T A95B

Qwen3.8-2.4T-A95B is a very large open-weights language model positioned for agentic and generative AI use cases. NVIDIA's developer blog identifies it as a 2.4T-parameter model with configurable reasoning, and the accompanying image filename frames it as part of the Qwen open-source line, signaling that the weights are publicly distributed rather than locked behind a proprietary API. The configurable reasoning behavior suggests developers can toggle the model's chain-of-thought style at inference time, which is useful when a workflow needs to trade latency for deeper deliberation on hard problems.

Beyond the headline scale, practical deployment guidance comes from both DigitalOcean and NVIDIA, which independently confirm hosting paths for the model. DigitalOcean's serverless Inference Engine listing makes the model accessible without infrastructure management, while NVIDIA's technical blog provides a serving recipe tuned for NVIDIA GB300 NVL72 systems, giving data-center operators a reference stack for running a 2.4T-parameter checkpoint at production scale. Categorized under NVIDIA's Agentic AI / Generative AI section, the model is aimed at teams building autonomous agents, retrieval-augmented assistants, and other generative applications that benefit from very large context handling and adjustable reasoning effort rather than from a small, fixed-cost chat endpoint.

Cortecsqwen3.8-2.4t-a95bqwen

Quick Info

Powered by
Provider
Cortecs
Model key
qwen3.8-2.4t-a95b
Release date
Aug 12, 2026
Last updated
Aug 12, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$2.50
Output token cost
$6.00

Limits

Output tokens
262,144 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen3.8 2.4T A95B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.8 2.4T A95B

Cortecs

CoverageBenchmark

NVIDIA's developer blog (August 12, 2026) documents day-0 serving support for Alibaba's open-weight Qwen3.8-2.4T-A95B on GB300 NVL72 hardware, explicitly naming the exact model variant. The post details the architecture as a 2.4-trillion-parameter fine-grained mixture-of-experts model with 95 billion activated paramete The blog outlines the full software ecosystem for self-hosting the open weights: NVIDIA NeMo AutoModel supports post-training via full supervised fine-tuning or memory-efficient LoRA on Hugging Face checkpoints, and open-source inference recipes are available for SGLang, vLLM, and NVIDIA Dynamo. A model-free NVIDIA NIM

Cortecs

CoverageBenchmark

Qwen3.8 2.4T A95B is an open-weight, MoE-based flagship designed for agentic workloads, with 2.4 trillion total parameters activating roughly 95 billion per token. According to the supplied excerpt, the model distributes its 2.4T parameters across 512 routed experts alongside a dedicated shared expert module, with the The model ships with a native 256K-token context window powered by a hybrid full and linear attention mechanism, and exposes configurable inference-time reasoning controls for dynamic trade-offs between depth and latency. Per the excerpt, it is immediately available on Lyceum Serverless Inference via an OpenAI-compatib

Cortecs

Coverage

Qwen3.8-2.4T-A95B was released on Hugging Face with a sophisticated chat template aimed at high-precision tool calling and integrated reasoning. Per the supplied excerpt, the model employs a strict XML-based format for function execution using tags such as tool_call, function, and parameter, with a ChatML-style syntax A key feature of this release is support for explicit reasoning instructions that can be processed before function calls, alongside a robust system prompt structure. According to the excerpt, function calls must follow a precise syntax with no suffixes to ensure compatibility with automated parsing systems, providing d

Cortecs

Coverage

Alibaba published the weights for Qwen3.8-2.4T-A95B on August 13, 2026, open-releasing its Max-tier flagship ten days after the API launch. Per the supplied excerpt, the model is a fine-grained MoE with 2.4T total parameters and 95B activated per token at roughly a 25:1 sparsity ratio, carrying 512 experts with 11 acti Thinking is required for all interactions, with reasoning effort selectable across low, medium, and xhigh (default), a maximum reasoning length of 262,144 tokens, and a maximum output of 131,072 tokens. The excerpt notes that the released checkpoint does not support multimodal input while the hosted Qwen3.8-Max does, a

Videos about Qwen3.8 2.4T A95B

More models around Qwen3.8 2.4T A95B