Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
SiliconFlow (China) logo

Model details

Qwen/Qwen3-VL-32B-Instruct

Qwen3-VL-32B-Instruct is a large-scale multimodal model built around 32 billion parameters that bridges visual perception with advanced textual reasoning. Its architecture introduces features like Interleaved-MRoPE and DeepStack, enabling the model to handle complex tasks such as long-context video understanding and precise spatial comprehension. This combination of capabilities positions the model as a practical choice for workflows that demand both image and video analysis alongside robust language reasoning, rather than just single-modality processing.

The model ships with an OpenAI-compatible API interface, which makes integration straightforward for developers already familiar with standard LLM tooling. Several infrastructure providers power deployments using vLLM, and the model supports fine-tuning through LoRA techniques, allowing teams to customize behavior on proprietary datasets. This adaptability makes Qwen3-VL-32B-Instruct suitable for organizations looking to deploy specialized vision-language pipelines without rebuilding from scratch, whether for document understanding, visual question answering, or interactive multimodal applications.

SiliconFlow (China)Qwen/Qwen3-VL-32B-Instructqwen

Quick Info

Powered by
Provider
SiliconFlow (China)
Model key
Qwen/Qwen3-VL-32B-Instruct
Release date
Oct 21, 2025
Last updated
Nov 25, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.20
Output token cost
$0.60

Limits

Output tokens
262,000 tokens
Context window
262,000 tokens

Transparent token rates

Compare Qwen/Qwen3-VL-32B-Instruct pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen/Qwen3-VL-32B-Instruct

SiliconFlow

Coverage

The official ModelScope model card for Qwen/Qwen3-VL-32B-Instruct (updated Oct 22, 2025) describes the model as the most powerful vision-language model in the Qwen series, available in a 33.36B parameter Image-Text-to-Text configuration under the Apache-2.0 license via the Transformers/Safetensors/PyTorch stack. It lis The card also documents architecture changes that distinguish this release: Interleaved-MRoPE for full-frequency positional allocation over time, width, and height to support long-horizon video reasoning; DeepStack, which fuses multi-level ViT features for finer-grained detail and sharper image-text alignment; and Text

SiliconFlow (China)

CoverageDiscourse

When loading Qwen/Qwen3-VL-32B-Instruct with vLLM, for example when using TRL’s GRPOTrainer with use_vllm=True, an error of the form AttributeError: 'Qwen3VLTextConfig' object has no attribute 'tie...

Videos about Qwen/Qwen3-VL-32B-Instruct

More models around Qwen/Qwen3-VL-32B-Instruct