Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
SiliconFlow (China) logo

Model details

Qwen/Qwen3-VL-30B-A3B-Instruct

Qwen3-VL-30B-A3B-Instruct represents a major step forward for the Qwen series as a vision-language model built around a Mixture of Experts architecture that activates only 3 billion parameters during inference despite a 30 billion parameter total footprint. This design delivers strong multimodal understanding while managing computational efficiency for real-world deployment. The model unifies text generation with visual comprehension across images and videos, excelling at perception of both real-world and synthetic content, 2D and 3D spatial grounding, and long-form visual reasoning. Its Visual Agent capabilities enable it to navigate PC and mobile interfaces, recognize interface elements, and complete multi-step tasks by invoking tools, while its Visual Coding Boost allows it to generate Draw.io diagrams and HTML/CSS/JS from visual inputs. Enhanced spatial perception lets it judge object positions, viewpoints, and occlusions with improved 2D grounding and emerging 3D reasoning for embodied AI tasks.

The Instruct variant optimizes the model for instruction-following across general multimodal tasks, making it well-suited for document AI, OCR workflows, UI assistance, spatial task automation, and agent research. Its expanded OCR capability now supports 32 languages with robustness in challenging conditions like low light, blur, and tilt, plus better handling of rare or ancient characters. The model handles hours-long video with full recall and second-level indexing thanks to its native 256K context window, expandable to 1M tokens, enabling comprehensive analysis of lengthy visual content. Performance benchmarks position it competitively against leading models in STEM reasoning, visual question answering, and multimodal agent tasks. The availability of LoRA fine-tuning on its MoE architecture opens pathways for domain adaptation, while the existence of a companion Thinking variant provides enhanced reasoning capabilities when deeperChain-of-thought processing is needed.

SiliconFlow (China)Qwen/Qwen3-VL-30B-A3B-Instructqwen

Quick Info

Powered by
Provider
SiliconFlow (China)
Model key
Qwen/Qwen3-VL-30B-A3B-Instruct
Release date
Oct 5, 2025
Last updated
Nov 25, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.29
Output token cost
$1.00

Limits

Output tokens
262,000 tokens
Context window
262,000 tokens

Transparent token rates

Compare Qwen/Qwen3-VL-30B-A3B-Instruct pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen/Qwen3-VL-30B-A3B-Instruct

SiliconFlow

Coverage

QuantTrio published an AWQ quantization of Qwen/Qwen3-VL-30B-A3B-Instruct on Hugging Face, explicitly naming the exact base model variant. The model card describes Qwen3-VL as a 30B-A3B MoE vision-language model with a native 256K context window expandable to 1M, Interleaved-MRoPE positional embeddings, DeepStack multi The card further enumerates Qwen3-VL capability upgrades including a Visual Agent for PC/mobile GUI operation, visual coding (Draw.io/HTML/CSS/JS generation from images), advanced spatial perception with 2D/3D grounding, enhanced multimodal reasoning for STEM/math, broader visual recognition coverage, and expanded OCR

Videos about Qwen/Qwen3-VL-30B-A3B-Instruct

More models around Qwen/Qwen3-VL-30B-A3B-Instruct