NEAR AI Cloud
vLLM Ascend officially validates Qwen3-VL-30B-A3B-Instruct on Huawei Ascend hardware in the v0.13.0 cycle, treating it as a sparse MoE vision-language model with about 30B total parameters and roughly 3B activated per token. The tutorial covers supported features, image and video online deployment, offline inference, functional verification, and accuracy and performance evaluation, with deployment guidance for Atlas 800 A3 and A2 nodes. Both BF16 and w8a8-mxfp8 quantized weights are supported, giving Ascend users a ready-to-run multimodal stack. The documentation walks through Docker image and source installation, then explains shared multi-node weight deployment paths so that distributed serving jobs can locate the same model directory. It links into Ascend-specific feature configuration and performance tuning for long-context multimodal workloads. This positions Qwen3-VL-30B-A3B-Instruct as a first-class Ascend target alongside mainstream vLLM-validated models for vision-language serving.
