Currently listed through these providers:
Model details
Qwen3 VL 30B A3B Instruct
The Qwen3 VL 30B A3B Instruct model is a powerful multimodal system designed to unify high-level text generation with sophisticated visual perception. By accepting inputs such as images, text, and bounding boxes, it enables a wide range of capabilities including multi-modal dialogue, image detection, and multi-image reasoning. The model is engineered to excel in complex environments, demonstrating particular strength in 2D and 3D spatial grounding, long-form visual comprehension, and the interpretation of both real-world and synthetic categories.
Built to support agentic workflows, the model is optimized for instruction-following across diverse tasks like video timeline alignment, GUI automation, and visual coding from sketches. Its architecture allows it to handle multi-turn instructions effectively, making it a strong candidate for document AI, OCR, and UI assistance. With performance that rivals flagship models in STEM, VQA, and reasoning benchmarks, it serves as a robust tool for developers looking to integrate advanced multimodal functions into their product roadmaps, from simple visual content recognition to complex, multi-file code generation.
Quick Info
Powered by- Provider
- DevPass (LLM Gateway)
- Model key
- qwen3-vl-30b-a3b-instruct
- Release date
- Oct 2, 2025
- Last updated
- Oct 2, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.15
- Output token cost
- $0.60
Limits
- Output tokens
- 8,192 tokens
- Context window
- 262,144 tokens
Transparent token rates
Compare Qwen3 VL 30B A3B Instruct pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.