Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
DevPass (LLM Gateway) logo

Model details

Qwen3 VL 30B A3B Instruct

The Qwen3 VL 30B A3B Instruct model is a powerful multimodal system designed to unify high-level text generation with sophisticated visual perception. By accepting inputs such as images, text, and bounding boxes, it enables a wide range of capabilities including multi-modal dialogue, image detection, and multi-image reasoning. The model is engineered to excel in complex environments, demonstrating particular strength in 2D and 3D spatial grounding, long-form visual comprehension, and the interpretation of both real-world and synthetic categories.

Built to support agentic workflows, the model is optimized for instruction-following across diverse tasks like video timeline alignment, GUI automation, and visual coding from sketches. Its architecture allows it to handle multi-turn instructions effectively, making it a strong candidate for document AI, OCR, and UI assistance. With performance that rivals flagship models in STEM, VQA, and reasoning benchmarks, it serves as a robust tool for developers looking to integrate advanced multimodal functions into their product roadmaps, from simple visual content recognition to complex, multi-file code generation.

DevPass (LLM Gateway)qwen3-vl-30b-a3b-instructqwen

Quick Info

Powered by
Provider
DevPass (LLM Gateway)
Model key
qwen3-vl-30b-a3b-instruct
Release date
Oct 2, 2025
Last updated
Oct 2, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$0.60

Limits

Output tokens
8,192 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen3 VL 30B A3B Instruct pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3 VL 30B A3B Instruct

Videos about Qwen3 VL 30B A3B Instruct

More models around Qwen3 VL 30B A3B Instruct