Currently listed through these providers:
Model details
Qwen3.5 397B-A17B
Qwen3.5-397B-A17B is a large-scale mixture-of-experts model in the Qwen3.5 family that is designed around four practical pillars: multimodal understanding, long-context inference, multi-token prediction speculative decoding, and W8A8 quantized deployment. The combination of an MoE architecture with W8A8 weight and activation quantization points to a design intent of retaining rich capability while keeping inference efficient enough for production traffic, and the explicit mention of long-context inference suggests the model is meant for workloads that need to reason over extended documents, codebases, or multimedia sequences rather than only short prompts.
Beyond its architecture, the model is positioned for real-world serving rather than purely research use. Official support in the vLLM Ascend stack begins with v0.17.0rc1 and continues through later versions, with documented validation paths for single-node online deployment, multi-node deployment, and Prefill-Decode disaggregation, along with accuracy and performance evaluation steps. The presence of a Qwen3.5-397B-A17B container listing on the NVIDIA NGC catalog under the NIM team further indicates that NVIDIA-hosted deployment artifacts are available, giving teams multiple inference paths. Overall, the model fits well for organizations that need a multimodal MoE system capable of handling long contexts with quantized, speculative-decoding-assisted serving on dedicated AI hardware.
Quick Info
Powered by- Provider
- TokenGo
- Model key
- qwen/qwen3.5-397b-a17b
- Release date
- Feb 15, 2026
- Last updated
- Feb 15, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.40
- Output token cost
- $2.65
Limits
- Output tokens
- 65,536 tokens
- Context window
- 262,144 tokens