Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
LLM Gateway logo

Model details

Qwen3.8 2.4T A95B (Together AI)

Qwen3.8-2.4T-A95B is the flagship of Qwen's 3.8 generation and the largest model the team has released, built around a sparse Mixture-of-Experts architecture with 2.4 trillion total parameters. It is also Qwen's first multimodal model above the one-trillion-parameter mark, designed to process text and images together. The architecture reflects a deliberate push toward flagship-scale capacity for workloads that benefit from broad world knowledge and deep reasoning, while the sparse MoE design keeps inference efficient relative to its overall size.

The model is aimed squarely at coding and long-horizon agentic work, with always-on thinking and adjustable reasoning effort that lets callers trade response speed against deliberation depth. A roughly one-million-token context window allows it to ingest entire repositories or thousands of pages in a single pass, supporting extended multi-step tasks without losing earlier context. Practical fit lies in complex software engineering, agent pipelines that need sustained planning, and multimodal reasoning scenarios where text and images must be interpreted together.

LLM Gatewaytogether-ai/qwen3.8-2.4t-a95bqwen

Quick Info

Powered by
Provider
LLM Gateway
Model key
together-ai/qwen3.8-2.4t-a95b
Release date
Aug 12, 2026
Last updated
Aug 12, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$2.00
Output token cost
$6.00

Limits

Output tokens
909,000 tokens
Context window
1,010,000 tokens

Transparent token rates

Compare Qwen3.8 2.4T A95B (Together AI) pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.8 2.4T A95B (Together AI)

LLM Gateway

CoverageBenchmark

Alibaba's Qwen3.8-Max, a 2.4-trillion-parameter mixture-of-experts model with 95 billion active parameters, first went live through Alibaba's QwenCloud API on August 2, 2026, with the open-weight variant packaged as Qwen3.8-2.4T-A95B following on Hugging Face and ModelScope on August 13, according to a September 7, 202 The recap situates the release in a crowded September for frontier AI, noting that OpenAI, Anthropic, and Google all shipped new models within days of each other and that Alibaba's decision to hand over full model weights has become a sharp example of open source pushing back on increasingly closed labs. It frames the

LLM Gateway

Coverage

Alibaba released the open weights for Qwen3.8-2.4T-A95B (also styled Qwen3.8-Max), a fine-grained mixture-of-experts model with 2.4 trillion total parameters and 95 billion activated per token, according to an August 12, 2026 NVIDIA Technical Blog post. The architecture pairs the MoE design with a hybrid of full and li Without additional tuning, the model achieves over 4,000 tokens per second per GPU and over 350 tokens per second per user on NVIDIA GB300 NVL72 systems in FP8 precision on day 0, with NVIDIA indicating further gains from NVFP4 optimizations. NVIDIA NeMo AutoModel supports post-training the model via full supervised fi

Videos about Qwen3.8 2.4T A95B (Together AI)

More models around Qwen3.8 2.4T A95B (Together AI)