Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Together AI logo

Model details

Qwen3.5 397B A17B

Qwen3.5-397B-A17B is a large-scale mixture-of-experts model in the Qwen3.5 family, configured with 397B total parameters and 17B active parameters per forward pass. The architecture is built to combine multimodal understanding with long-context inference, making it suitable for deployments that need to process extended documents alongside visual inputs. Its MoE design aims to balance broad capacity against the per-token compute cost, keeping active inference tractable while still exposing a very large parameter pool for reasoning and tool use.

In production environments the model is validated for deployment with vLLM-Ascend on Ascend hardware, supporting BF16 and W8A8 quantization, chunked prefill, automatic prefix caching, speculative decoding with MTP, asynchronous scheduling, tensor parallelism, and expert parallelism. It is also packaged as an NVIDIA NIM container, giving teams a second supported inference path for enterprise serving. The model has been the subject of recent Together AI platform improvements to fine-tuning training quality, indicating ongoing investment in tuning workflows for organizations adapting it to domain-specific tasks.

Together AIQwen/Qwen3.5-397B-A17Bqwendeprecated

Quick Info

Powered by
Provider
Together AI
Model key
Qwen/Qwen3.5-397B-A17B
Release date
Feb 16, 2026
Last updated
Jun 15, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.60
Output token cost
$3.60

Limits

Output tokens
130,000 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen3.5 397B A17B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.5 397B A17B

Mixlayer

CoverageBenchmark

An independent technical analysis of Qwen3.5-397B-A17B published on InferenceX confirms the model as the first open-weights release in Alibaba's Qwen3.5 series, explicitly naming the variant with 397B total parameters and 17B activated per forward pass. The page documents the February 2026 announcement by the Qwen team The same InferenceX entry details that Qwen3.5-397B-A17B unifies thinking and non-thinking behavior in a single checkpoint, operating in thinking mode by default with `...` tags before the answer and dropping the `/think` and `/nothink` soft switches used in Qwen3, so non-thinking responses are controlled via API param

Mixlayer

Coverage

According to the Baidu encyclopedia entry, Alibaba launched Qwen3.5 on February 16, 2026, as a flagship large language model series comprising Qwen3.5-Plus and Qwen3.5-397B-A17B, with the latter explicitly named as a 397-billion-parameter MoE that activates only 17 billion parameters per token and is reported to reduce The same entry reports that the hosted Qwen3.5-Plus API was priced at 0.8 RMB per million tokens, scored 87.8 on the MMLU-Pro knowledge reasoning benchmark, and that at 32K context length the model achieves an 8.6x increase in inference throughput. It also describes Day-0 ecosystem adaptation by NVIDIA, AMD, and Apple

Together AI

Official sourceRelease Notes

The Together AI changelog (dated August 27, 2026) confirms under August 24, 2026 "Improvements" that fine-tuning training quality has been improved for Qwen/Qwen3.5-397B-A17B, alongside other Qwen3.5 variants (0.8B through 122B-A10B) and Qwen3.6 models. This is first-party evidence of a recent, directly relevant platfo The same changelog entry contextualizes the broader platform state, including August 27 model deprecations (Nemotron-3-ultra-550b-a55b, DeepSeek-V4-Pro in favor of DeepSeek-V4-Pro-0813, Kimi-K2.7-Code, gemma-4-31b-it), August 26 batch-jobs CLI support, and the August 26 serverless release of zai-org/GLM-5.3-Flash at $0

Videos about Qwen3.5 397B A17B

More models around Qwen3.5 397B A17B