Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Hugging Face logo

Model details

Qwen3.5-397B-A17B

Alibaba’s Qwen3.5-397B-A17B is a native vision-language foundation model designed for reasoning, coding, multimodal understanding, and agent-based work. It uses the Qwen3-Next hybrid design, combining Gated Delta Network linear attention with high-sparsity sparse mixture-of-experts routing, and activates 17 billion of its 397 billion parameters for each forward pass. Its 60-layer architecture also includes gated full attention, multi-token prediction, and early fusion of visual and textual information.

The model is best suited to demanding workflows such as repository-level software tasks, terminal and web-research agents, tool use, mathematical and scientific reasoning, and joint text-image analysis. Qwen-reported evaluations include 76.4 on SWE-bench Verified, 83.6 on LiveCodeBench v6, 88.4 on GPQA Diamond, 87.8 on MMLU-Pro, and 85.0 on MMMU; these results indicate broad technical strength, although they come from vendor-specific harnesses. Its native long-context design and published Apache 2.0 checkpoint make it relevant for teams building capable open-weight multimodal agents.

Hugging FaceQwen/Qwen3.5-397B-A17Bqwen

Quick Info

Powered by
Provider
Hugging Face
Model key
Qwen/Qwen3.5-397B-A17B
Release date
Feb 1, 2026
Last updated
Feb 1, 2026
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.60
Output token cost
$3.60

Limits

Output tokens
32,768 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen3.5-397B-A17B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.5-397B-A17B

Nebius Token Factory

CoverageBenchmark

SemiAnalysis's InferenceX page characterizes Qwen3.5-397B-A17B as the first open-weights model in Alibaba's Qwen3.5 series, citing the Alibaba Cloud Qwen team announcement and the Hugging Face model card as primary sources. It reports a February 2026 release window (February 16 via OpenRouter, February 17 via Artificia The same analysis describes Qwen3.5-397B-A17B as an explicitly agentic and multimodal foundation model targeting reasoning, coding, agent capabilities, and multimodal understanding, and notes that early-fusion multimodal training is reported to achieve cross-generational parity with Qwen3 while outperforming the Qwen3-

Nebius Token Factory

Coverage

Alibaba announced on February 16, 2026 that it has open-sourced Qwen3.5, beginning with the Qwen3.5-397B-A17B checkpoint, which Alibaba also refers to as Qwen3.5-Plus. The post describes it as a natively multimodal foundation model trained across text, images, and video, with support extended to 201 languages and diale The same announcement highlights the model's inference-efficiency design, combining a linear attention mechanism with a sparse mixture-of-experts (MoE) architecture to lower compute requirements. According to the post, Qwen3.5-397B-A17B demonstrates strong performance on language understanding and reasoning, code gener

Nebius Token Factory

Coverage

The Baidu Baike entry on Qwen3.5 confirms that Alibaba's Qwen3.5 series, launched on February 16, 2026, comprises two versions — Qwen3.5-Plus and Qwen3.5-397B-A17B — built around a Hybrid Attention Mechanism and a Sparse Mixture-of-Experts architecture optimized for logical reasoning, mathematical computation, and code The same entry traces development milestones in early 2026 — Alibaba's January announcement of the Qwen3.5 plan, a February 9 code merge request appearing on the Hugging Face open-source project page, and showcase events on February 11 — and reports a score of 87.8 on the MMLU-Pro knowledge reasoning benchmark for the

Videos about Qwen3.5-397B-A17B

More models around Qwen3.5-397B-A17B