SiliconFlow
The official ModelScope model card for Qwen/Qwen3-VL-32B-Instruct (updated Oct 22, 2025) describes the model as the most powerful vision-language model in the Qwen series, available in a 33.36B parameter Image-Text-to-Text configuration under the Apache-2.0 license via the Transformers/Safetensors/PyTorch stack. It lis The card also documents architecture changes that distinguish this release: Interleaved-MRoPE for full-frequency positional allocation over time, width, and height to support long-horizon video reasoning; DeepStack, which fuses multi-level ViT features for finer-grained detail and sharper image-text alignment; and Text