Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
LLM Gateway logo

Model details

Qwen3.8 Max (Alibaba Cloud)

Qwen3.8 Max is a large multimodal model from Alibaba Cloud's Qwen family, designed to handle text, images, video, and PDF inputs while producing text outputs. It is positioned as a flagship chat model with first-class reasoning tokens and tool-calling capabilities, making it suitable for agentic workflows that combine grounded perception with chain-of-thought deliberation. The release is tied to a Hugging Face model card for a configuration labeled Qwen3.8-2.4T-A95B, which signals a Mixture-of-Experts lineage with an extremely large total parameter budget, even though the active per-token parameter count is much smaller.

Independent benchmark tracking gives Qwen3.8 Max an overall decision score of 79.2 out of 100, placing it sixth among 226 tracked models as of late August 2026. Its strongest published category is Multimodal and Grounded, where it ranks first, suggesting particular strength at screenshots, documents, charts, and other visually anchored tasks. A one-the cataloged API limit and reasoning-token support let it work over very long documents, while Apache-2.0 licensing and open weights make self-hosting and fine-tuning practical for teams that need full control over their stack.

LLM Gatewayalibaba/qwen3.8-maxqwen

Quick Info

Powered by
Provider
LLM Gateway
Model key
alibaba/qwen3.8-max
Release date
Aug 3, 2026
Last updated
Aug 3, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$2.00
Output token cost
$6.00

Limits

Output tokens
131,072 tokens
Context window
983,616 tokens

Transparent token rates

Compare Qwen3.8 Max (Alibaba Cloud) pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.8 Max (Alibaba Cloud)

LLM Gateway

CoverageBenchmark

BenchLM.ai maintains a dedicated model record for Qwen3.8 Max (Alibaba) that reports a release date of August 3, 2026, an open-weight designation, a 1M-token context window, and an architecture tag of Qwen3.8-2.4T-A95B consistent with the 2.4T parameter figure seen on the LVBench leaderboard. The page lists a composite Speed measurements on the page include 38 tokens/second sustained throughput and a 55.38s first-token latency, while Reasoning, Math, and Multilingual categories are marked as not rank-eligible on this evaluator. As an aggregator, BenchLM.ai computes its own composite scores from heterogeneous benchmarks, so the percen

LLM Gateway

Coverage

A third-party event page from observatory.srmdn.com (dated August 3, 2026) reports that Alibaba Cloud has introduced Qwen3.8-Max, positioning it as a flagship model aimed at complex reasoning, coding, multimodal tasks, and long-horizon workflows, with access provided through the Qwen API. The excerpt frames the announc The event page supplies no concrete technical specifications in the excerpt (no parameter counts, context window lengths, benchmark figures, or pricing), and no direct citation to an Alibaba or Qwen Team primary source is visible. The source host (observatory.srmdn.com) is a third-party observatory, and its capability

Videos about Qwen3.8 Max (Alibaba Cloud)

More models around Qwen3.8 Max (Alibaba Cloud)