Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Mixlayer logo

Model details

Qwen3.5 397B A17B

Qwen3.5-397B-A17B sits within the broader Qwen3.5 family as a large-scale Mixture-of-Experts (MoE) model, a design choice that typically allows very large total capacity while activating only a fraction of the parameters per token for efficient inference. According to the vLLM Ascend deployment documentation, the model brings together multimodal capability, long-context inference, MTP (Multi-Token Prediction) speculative decoding, and W8A8 weight-and-activation quantization, signalling an emphasis on high-throughput, production-grade serving rather than purely research-scale experimentation. The "A17B" segment of the name strongly suggests an active-parameter footprint of roughly 17 billion within a total of around 397 billion parameters, though that specific breakdown is implied by the naming convention rather than spelled out in the cited sources.

The model has an NGC container entry under the nim/qwen organization path, which points to coordinated packaging alongside NVIDIA's NIM stack, while its first-class support in vllm-ascend from v0.17.0rc1 onward indicates that operators can run it with features such as BF16 and W8A8 quantization, chunked prefill, automatic prefix caching, speculative decoding, asynchronous scheduling, tensor parallelism, and expert parallelism. This combination makes it well suited to workloads that need to ingest very long documents, generate extended outputs, or take advantage of speculative decoding for lower latency, particularly on Ascend-accelerated infrastructure. Its open-weight availability further broadens its appeal for teams that want to self-host, fine-tune, or integrate the model into private pipelines without depending on a closed API.

Mixlayerqwen/qwen3.5-397b-a17bqwen

Quick Info

Powered by
Provider
Mixlayer
Model key
qwen/qwen3.5-397b-a17b
Release date
Mar 18, 2026
Last updated
Mar 18, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.60
Output token cost
$3.60

Limits

Output tokens
262,144 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen3.5 397B A17B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.5 397B A17B

Mixlayer

CoverageBenchmark

An independent technical analysis of Qwen3.5-397B-A17B published on InferenceX confirms the model as the first open-weights release in Alibaba's Qwen3.5 series, explicitly naming the variant with 397B total parameters and 17B activated per forward pass. The page documents the February 2026 announcement by the Qwen team The same InferenceX entry details that Qwen3.5-397B-A17B unifies thinking and non-thinking behavior in a single checkpoint, operating in thinking mode by default with `...` tags before the answer and dropping the `/think` and `/nothink` soft switches used in Qwen3, so non-thinking responses are controlled via API param

Mixlayer

Coverage

According to the Baidu encyclopedia entry, Alibaba launched Qwen3.5 on February 16, 2026, as a flagship large language model series comprising Qwen3.5-Plus and Qwen3.5-397B-A17B, with the latter explicitly named as a 397-billion-parameter MoE that activates only 17 billion parameters per token and is reported to reduce The same entry reports that the hosted Qwen3.5-Plus API was priced at 0.8 RMB per million tokens, scored 87.8 on the MMLU-Pro knowledge reasoning benchmark, and that at 32K context length the model achieves an 8.6x increase in inference throughput. It also describes Day-0 ecosystem adaptation by NVIDIA, AMD, and Apple

Videos about Qwen3.5 397B A17B

More models around Qwen3.5 397B A17B