Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
SiliconFlow (China) logo

Model details

Qwen/Qwen3-VL-30B-A3B-Thinking

The Qwen3-VL-30B-A3B-Thinking model is a multimodal system built on a Mixture-of-Experts architecture, featuring 30 billion total parameters with 3 billion active during inference. It utilizes a hybrid multimodal block that merges a visual encoder with a language core, connected by an Interleaved-MRoPE positioning mechanism. This design allows the model to precisely distribute frequency features across time and space, enabling stable 3D grounding and efficient synchronization of video timestamps with text. By treating visual representations as n-dimensional tokens within a unified reasoning context, the model achieves high-fidelity comprehension of complex visual data, from GUI elements to intricate spatial relationships.

Designed for high-level cognitive tasks, the model features a specialized reasoning mode that processes information step-by-step to provide evidence-based, causal conclusions. This capability is supported by a robust training lineage that emphasizes STEM and mathematical logic, making it well-suited for demanding applications like visual coding, scientific analysis, and long-form document parsing. With native support for a 256K-token context window that can scale up to 1M tokens, the model maintains data coherence across multi-hour videos and extensive text, positioning it as a versatile tool for agentic workflows, automated UI interaction, and advanced multimodal research.

SiliconFlow (China)Qwen/Qwen3-VL-30B-A3B-Thinkingqwen

Quick Info

Powered by
Provider
SiliconFlow (China)
Model key
Qwen/Qwen3-VL-30B-A3B-Thinking
Release date
Oct 11, 2025
Last updated
Nov 25, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.29
Output token cost
$1.00

Limits

Output tokens
262,000 tokens
Context window
262,000 tokens

Transparent token rates

Compare Qwen/Qwen3-VL-30B-A3B-Thinking pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen/Qwen3-VL-30B-A3B-Thinking

SiliconFlow

Coverage

The ModelScope model card for Qwen/Qwen3-VL-30B-A3B-Thinking lists it as an Image-Text-to-Text model in the Qwen3-VL family, tagged qwen3_vl_moe, with 31.07B parameters, Transformers/Safetensors/PyTorch weights, and an apache-2.0 license. The card reports ~47,889 downloads, a 62.15GB size, and a last-updated stamp of N Architecture and capability details on the card describe Interleaved-MRoPE for full-frequency time/width/height positional encoding, DeepStack multi-level ViT feature fusion, and Text-Timestamp Alignment that moves beyond T-RoPE for timestamp-grounded video reasoning. Stated capabilities include a Visual Agent that ope

Videos about Qwen/Qwen3-VL-30B-A3B-Thinking

More models around Qwen/Qwen3-VL-30B-A3B-Thinking