Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
NovitaAI logo

Model details

qwen/qwen3-vl-30b-a3b-thinking

Qwen3-VL-30B-A3B-Thinking is the reasoning-enhanced vision-language variant within the Qwen3 VL family, designed to unify strong text generation with deep visual understanding across images and videos. Its architecture builds on the core Qwen3 VL foundation but adds an extended thinking mechanism that strengthens performance on complex reasoning tasks—particularly in STEM fields and mathematical problem-solving. The model excels at perceiving real-world and synthetic visual categories, performing precise 2D and 3D spatial grounding, and sustaining long-form visual comprehension that supports applications ranging from document AI and OCR to GUI automation and spatial reasoning for embodied AI systems. The vision encoder and language model components are available separately in GGUF format, enabling flexible deployment across CPU, NVIDIA CUDA, Apple Silicon Metal, and other GGUF-compatible backends for developers who want to run the model on personal devices or custom hardware.

The Thinking edition reflects a deliberate post-training strategy that cultivates reasoning capabilities beyond standard instruction tuning, positioning this variant for tasks where step-by-step logic matters alongside multimodal perception. Text generation performance is engineered to match flagship Qwen3 text models, which means the model carries forward advances in language quality from the broader Qwen3 lineup into visual contexts. For agentic workflows, the model handles multi-image multi-turn instructions, aligns video timelines, operates PC and mobile GUIs by recognizing interface elements and invoking tools, and supports visual coding workflows that can generate, debug, and refine UI code from sketches or screenshots. This combination of visual grounding, extended context, and reasoning-enhanced output makes it well-suited for research environments focused on multimodal agents, spatial understanding, and autonomous task completion.

NovitaAIqwen/qwen3-vl-30b-a3b-thinking

Quick Info

Powered by
Provider
NovitaAI
Model key
qwen/qwen3-vl-30b-a3b-thinking
Release date
Oct 11, 2025
Last updated
Oct 11, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.20
Output token cost
$1.00

Limits

Output tokens
32,768 tokens
Context window
131,072 tokens

Latest news about qwen/qwen3-vl-30b-a3b-thinking

NovitaAI

CoverageBenchmark

Qwen3 VL 30B A3B Thinking pricing: $0.13/M input, $0.60/M output. Compare with 10 similar models, see benchmarks, and find the cheapest provider.

SiliconFlow

Coverage

The ModelScope model card for Qwen/Qwen3-VL-30B-A3B-Thinking lists it as an Image-Text-to-Text model in the Qwen3-VL family, tagged qwen3_vl_moe, with 31.07B parameters, Transformers/Safetensors/PyTorch weights, and an apache-2.0 license. The card reports ~47,889 downloads, a 62.15GB size, and a last-updated stamp of N Architecture and capability details on the card describe Interleaved-MRoPE for full-frequency time/width/height positional encoding, DeepStack multi-level ViT feature fusion, and Text-Timestamp Alignment that moves beyond T-RoPE for timestamp-grounded video reasoning. Stated capabilities include a Visual Agent that ope

Videos about qwen/qwen3-vl-30b-a3b-thinking