Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Merge Gateway logo

Model details

Qwen3-VL 30B-A3B Instruct

We haven't written an overview of this model yet. New models can take a few days to gather enough reliable coverage, so check back soon.

Merge Gatewayqwen/qwen3-vl-30b-a3b-instructqwen

Quick Info

Powered by
Provider
Merge Gateway
Model key
qwen/qwen3-vl-30b-a3b-instruct
Release date
Oct 4, 2025
Last updated
Oct 4, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.108
Output token cost
$0.43

Limits

Output tokens
32,768 tokens
Context window
131,072 tokens

Transparent token rates

Compare Qwen3-VL 30B-A3B Instruct pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3-VL 30B-A3B Instruct

NEAR AI Cloud

CoverageDocumentation

vLLM Ascend officially validates Qwen3-VL-30B-A3B-Instruct on Huawei Ascend hardware in the v0.13.0 cycle, treating it as a sparse MoE vision-language model with about 30B total parameters and roughly 3B activated per token. The tutorial covers supported features, image and video online deployment, offline inference, functional verification, and accuracy and performance evaluation, with deployment guidance for Atlas 800 A3 and A2 nodes. Both BF16 and w8a8-mxfp8 quantized weights are supported, giving Ascend users a ready-to-run multimodal stack. The documentation walks through Docker image and source installation, then explains shared multi-node weight deployment paths so that distributed serving jobs can locate the same model directory. It links into Ascend-specific feature configuration and performance tuning for long-context multimodal workloads. This positions Qwen3-VL-30B-A3B-Instruct as a first-class Ascend target alongside mainstream vLLM-validated models for vision-language serving.

NEAR AI Cloud

Coverage

Qwen3-VL-30B-A3B-Instruct is the official weight release of the Qwen3-VL generation, described on the model card as the most powerful vision-language model in the series to date, available in both Dense and MoE variants with Instruct and Thinking editions. Key upgrades include a Visual Agent that operates PC and mobile GUIs, visual coding that generates Draw.io, HTML, CSS, and JS from images or video, and stronger 2D and 3D spatial grounding for embodied tasks. It expands OCR to 32 languages from 19, handles rare scripts and jargon, and keeps text understanding on par with pure LLMs. Architecturally, the model introduces Interleaved-MRoPE over time, width, and height for long-horizon video reasoning, DeepStack fusion of multi-level ViT features for finer image-text alignment, and Text-Timestamp Alignment that replaces T-RoPE for precise event localization. It supports a native 256K context expandable to 1M, enabling hours-long video comprehension with second-level indexing, and excels at multimodal reasoning across STEM and math, outputting evidence-based conclusions for complex questions.

NEAR AI Cloud

Coverage

The Hugging Face model card for QuantTrio/Qwen3-VL-30B-A3B-Instruct-AWQ explicitly names the subject model Qwen/Qwen3-VL-30B-A3B-Instruct as the base, providing an AWQ quantization tailored for efficient deployment. The card documents a 17 GB model file (updated 2025-10-04) and supplies a vLLM startup command along wit On the architecture side, the card details Interleaved-MRoPE for full-frequency time/width/height positional allocation, DeepStack multi-level ViT feature fusion for fine-grained image-text alignment, and Text-Timestamp Alignment replacing T-RoPE for timestamp-grounded video reasoning. These technical specifics are tie

Videos about Qwen3-VL 30B-A3B Instruct

More models around Qwen3-VL 30B-A3B Instruct