Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
SiliconFlow logo

Model details

Qwen/Qwen3-VL-30B-A3B-Thinking

Qwen3-VL-30B-A3B-Thinking is a multimodal model built on a Mixture-of-Experts architecture, utilizing 30 billion total parameters with 3 billion active parameters to balance performance and efficiency. Designed as a reasoning-enhanced variant within the Qwen series, it integrates advanced visual perception with high-level text generation. The model is engineered to excel in complex tasks such as STEM and mathematical analysis, while providing robust support for 2D and 3D spatial grounding. Its design intent focuses on creating a unified system capable of seamless text-vision fusion, allowing it to handle intricate visual inputs alongside long-form document and video comprehension.

The model leverages a specialized training lineage that emphasizes reasoning capabilities and agentic interaction, enabling it to operate PC and mobile GUIs, invoke tools, and translate visual sketches into functional code. By incorporating enhanced spatial and video dynamics comprehension, the model achieves high-quality recognition across diverse categories, including landmarks, flora, fauna, and rare characters. Its architecture supports native long-context processing, making it well-suited for analyzing hours-long video content and extensive documents. This combination of visual agent functionality and deep logical reasoning positions the model as a versatile tool for developers building applications that require precise visual grounding and evidence-based decision-making.

SiliconFlowQwen/Qwen3-VL-30B-A3B-Thinkingqwen

Quick Info

Powered by
Provider
SiliconFlow
Model key
Qwen/Qwen3-VL-30B-A3B-Thinking
Release date
Oct 11, 2025
Last updated
Nov 25, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.29
Output token cost
$1.00

Limits

Output tokens
262,000 tokens
Context window
262,000 tokens

Transparent token rates

Compare Qwen/Qwen3-VL-30B-A3B-Thinking pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen/Qwen3-VL-30B-A3B-Thinking

SiliconFlow

Coverage

The ModelScope model card for Qwen/Qwen3-VL-30B-A3B-Thinking lists it as an Image-Text-to-Text model in the Qwen3-VL family, tagged qwen3_vl_moe, with 31.07B parameters, Transformers/Safetensors/PyTorch weights, and an apache-2.0 license. The card reports ~47,889 downloads, a 62.15GB size, and a last-updated stamp of N Architecture and capability details on the card describe Interleaved-MRoPE for full-frequency time/width/height positional encoding, DeepStack multi-level ViT feature fusion, and Text-Timestamp Alignment that moves beyond T-RoPE for timestamp-grounded video reasoning. Stated capabilities include a Visual Agent that ope

SiliconFlow

CoverageBenchmark

Qwen3 VL 30B A3B Thinking pricing: $0.13/M input, $0.60/M output. Compare with 10 similar models, see benchmarks, and find the cheapest provider.

Videos about Qwen/Qwen3-VL-30B-A3B-Thinking

More models around Qwen/Qwen3-VL-30B-A3B-Thinking