Sulat.com
AI models
Qiniu logo

Model details

Qwen3.5 397B A17B

Large-scale mixture-of-experts design with multimodal input and long-context support characterizes this model, which combines 397B total parameters with an active 17B expert configuration per token. The architecture is built to deliver production-grade inference, integrating multimodal capability, long-context handling, multi-token prediction speculative decoding, and W8A8 quantized deployment paths for efficient serving. This design choice lets the model balance broad knowledge capacity against the per-token compute needed for responsive responses, making it well suited for complex reasoning workloads across text and image inputs.

Practical deployment is a central focus, with validation covering single-node and multi-node online deployment, Prefill-Decode disaggregation, and functional, accuracy, and performance evaluation pipelines. Supported feature matrices include BF16 and W8A8 quantization, chunked prefill, automatic prefix caching, speculative decoding, asynchronous scheduling, tensor parallelism, and expert-level parallelism for distributed serving. The model is optimized for Ascend hardware starting from vllm-ascend v0.17.0rc1, with Ascend95DT support from v0.23.0rc1, making it a strong fit for teams building scalable, low-latency inference systems that require multimodal inputs, extended context windows, and structured reasoning outputs.

Qiniuqwen3.5-397b-a17b

Quick Info

Powered by
Provider
Qiniu
Model key
qwen3.5-397b-a17b
Release date
Feb 22, 2026
Last updated
Feb 22, 2026
Input modalities
Output modalities
Capabilities

Limits

Output tokens
64,000 tokens
Context window
256,000 tokens

Latest news about Qwen3.5 397B A17B

Qiniu

CoverageBenchmark

Benchmark Qwen3.5 397B A17B API latency, throughput, and cost efficiency. Compare response speed, token output, and pricing for scalable AI workloads.

Qiniu

CoverageBenchmark

OpenRouter's listing for Qwen3.5 397B A17B confirms a release date of February 16, 2026, a 262K context window, and multimodal support spanning language, code generation, agent tasks, image understanding, video understanding, and GUI interactions. The listed base price is $0.39 per 1M input tokens and $2.34 per 1M outp Provider-level performance data from OpenRouter shows meaningful divergence: Alibaba Cloud Int. leads on throughput at 80 tps with 0.80s latency and 99.99% uptime, while DeepInfra offers the lowest latency at 0.53s and Parasail reaches 0.49s. Cache-read pricing varies widely, from $0.111/M (DigitalOcean) up to $0.55/M

Videos about Qwen3.5 397B A17B