SiliconFlow (China)
Benchmark Qwen3.5 397B A17B API latency, throughput, and cost efficiency. Compare response speed, token output, and pricing for scalable AI workloads.
Model details
Qwen3.5-397B-A17B is a large-scale open-weight foundation model built on a sparse Mixture-of-Experts architecture, giving it 397 billion total parameters while activating only a subset of them during any given forward pass. This dynamic parameter activation means developers can tap into large-model intelligence without the full computational burden of a dense 400B model at every step. The model pairs this MoE design with a hybrid architecture that blends linear attention mechanisms, enabling high-throughput inference with reduced latency and cost overhead. It functions as a native vision-language model, capable of processing text, images, and video inputs within a unified framework that was trained using early fusion techniques across multimodal tokens, achieving cross-generational parity with the standalone Qwen3 series.
The development lineage traces back to Qwen's emphasis on integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and accessibility for developers and enterprises. Qwen3.5 draws on post-training work that has produced artifacts compatible with Hugging Face Transformers, vLLM, SGLang, and other inference engines, making it practical to deploy across varied infrastructure. Its design specifically targets complex AI workflows involving reasoning, coding, and multimodal tasks, where its open-weight nature and efficient hybrid design make it viable for teams that want the power of a frontier-scale model without the typical deployment constraints.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
SiliconFlow (China)
Benchmark Qwen3.5 397B A17B API latency, throughput, and cost efficiency. Compare response speed, token output, and pricing for scalable AI workloads.