Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
SiliconFlow logo

Model details

Qwen/Qwen3-VL-32B-Instruct

Qwen3-VL-32B-Instruct is the Instruct edition of the Qwen3-VL family, positioned by its maintainers as a dense vision-language model that fuses text understanding with deeper visual perception and reasoning. The official Qwen repository describes it as supporting visual agent behavior on PC and mobile GUIs, generating code and diagrams from images, judging object positions and viewpoints, and expanding OCR coverage to 32 languages with robustness in low-light, blurred, or tilted conditions. It is presented as on par with pure LLMs on text tasks, with seamless text-vision fusion for unified comprehension across modalities.

Third-party listings characterize the model as a 32-billion-parameter multimodal system with multimodal fusion delivered through Interleaved-MRoPE and DeepStack architectures, a native 256K context expandable to 1M, and the ability to handle long documents and hours-long video with second-level indexing. Practical strengths highlighted across sources include stronger multimodal reasoning for STEM and math, broader visual recognition across celebrities, landmarks, flora and fauna, and improved long-document structure parsing. Independent evaluation work has started probing its limits, such as the Enginuity benchmark for engineering diagrams, suggesting it is increasingly being stress-tested on domain-specific technical imagery rather than only general visual question answering.

SiliconFlowQwen/Qwen3-VL-32B-Instructqwen

Quick Info

Powered by
Provider
SiliconFlow
Model key
Qwen/Qwen3-VL-32B-Instruct
Release date
Oct 21, 2025
Last updated
Nov 25, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.20
Output token cost
$0.60

Limits

Output tokens
262,000 tokens
Context window
262,000 tokens

Transparent token rates

Compare Qwen/Qwen3-VL-32B-Instruct pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen/Qwen3-VL-32B-Instruct

SiliconFlow

CoverageBenchmark

The Enginuity paper (arXiv:2606.03410, June 2026) introduces the first open benchmark for evaluating vision-language models on engineering diagrams such as exploded parts views, electrical schematics, hydraulic circuits, and assembly drawings, drawn from U.S. military service and repair manuals. It defines two tasks: s The authors also find that token-overlap metrics under-report model capability on technical descriptions by roughly 2–6× relative to semantic similarity, motivating LLM-as-judge calibration for domain-specific evaluation of VLMs such as Qwen3-VL-32B-Instruct. They release the dataset, expert annotations, evaluation har

SiliconFlow

Coverage

The official ModelScope model card for Qwen/Qwen3-VL-32B-Instruct (updated Oct 22, 2025) describes the model as the most powerful vision-language model in the Qwen series, available in a 33.36B parameter Image-Text-to-Text configuration under the Apache-2.0 license via the Transformers/Safetensors/PyTorch stack. It lis The card also documents architecture changes that distinguish this release: Interleaved-MRoPE for full-frequency positional allocation over time, width, and height to support long-horizon video reasoning; DeepStack, which fuses multi-level ViT features for finer-grained detail and sharper image-text alignment; and Text

SiliconFlow

CoverageBenchmark

The OpenRouter model card describes Qwen3-VL-32B-Instruct as a 32-billion-parameter multimodal vision-language model with text, image, and video input/output modalities, a 131K context window, OCR support across 32 languages, and multimodal fusion via Interleaved-MRoPE and DeepStack architectures. It lists a release da Importantly for this subject, the OpenRouter listing shows a single hosting provider — Alibaba Cloud Int. — at 100% token share, with no SiliconFlow provider entry. This means the page corroborates the underlying model's general specifications, pricing tier, and performance characteristics, but does not independently c

SiliconFlow

CoverageBenchmark

Qwen3 VL 32B Instruct pricing: $0.10/M input, $0.42/M output. Compare with 10 similar models, see benchmarks, and find the cheapest provider.

SiliconFlow

CoverageDiscourse

When loading Qwen/Qwen3-VL-32B-Instruct with vLLM, for example when using TRL’s GRPOTrainer with use_vllm=True, an error of the form AttributeError: 'Qwen3VLTextConfig' object has no attribute 'tie...

Videos about Qwen/Qwen3-VL-32B-Instruct

More models around Qwen/Qwen3-VL-32B-Instruct