Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
SiliconFlow logo

Model details

Qwen/Qwen3-VL-30B-A3B-Instruct

Qwen3-VL-30B-A3B-Instruct serves as a versatile vision-language model designed to unify high-level text generation with deep visual perception. Built to handle both images and videos, the model excels at tasks requiring spatial reasoning, such as judging object positions and viewpoints, while supporting 2D and 3D grounding. Its architecture is engineered for agentic workflows, allowing it to operate PC and mobile graphical user interfaces by recognizing elements and invoking tools. Beyond visual tasks, the model provides robust support for visual coding, enabling the generation of functional UI components from sketches, and offers expanded OCR capabilities that handle diverse languages, rare characters, and complex document structures.

The model benefits from a broad, high-quality pretraining process that enables it to recognize a wide array of real-world and synthetic categories, from landmarks to specialized jargon. By integrating seamless text-vision fusion, it achieves text understanding performance comparable to pure large language models, making it a strong candidate for STEM, mathematics, and logical analysis. Designed for flexibility, the model supports long-context processing, capable of indexing hours of video and extensive documents with second-level precision. These strengths position it as an effective tool for developers building applications that require sophisticated multimodal reasoning, automated visual assistance, and reliable, evidence-based instruction following.

SiliconFlowQwen/Qwen3-VL-30B-A3B-Instructqwen

Quick Info

Powered by
Provider
SiliconFlow
Model key
Qwen/Qwen3-VL-30B-A3B-Instruct
Release date
Oct 5, 2025
Last updated
Nov 25, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.29
Output token cost
$1.00

Limits

Output tokens
262,000 tokens
Context window
262,000 tokens

Transparent token rates

Compare Qwen/Qwen3-VL-30B-A3B-Instruct pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen/Qwen3-VL-30B-A3B-Instruct

SiliconFlow

Coverage

QuantTrio published an AWQ quantization of Qwen/Qwen3-VL-30B-A3B-Instruct on Hugging Face, explicitly naming the exact base model variant. The model card describes Qwen3-VL as a 30B-A3B MoE vision-language model with a native 256K context window expandable to 1M, Interleaved-MRoPE positional embeddings, DeepStack multi The card further enumerates Qwen3-VL capability upgrades including a Visual Agent for PC/mobile GUI operation, visual coding (Draw.io/HTML/CSS/JS generation from images), advanced spatial perception with 2D/3D grounding, enhanced multimodal reasoning for STEM/math, broader visual recognition coverage, and expanded OCR

Videos about Qwen/Qwen3-VL-30B-A3B-Instruct

More models around Qwen/Qwen3-VL-30B-A3B-Instruct