Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Kilo Gateway logo

Model details

Qwen3 VL 235B A22B Instruct

Qwen3 VL 235B A22B Instruct is a flagship multimodal vision-language model released in September 2025 by the Qwen team at Alibaba Cloud, built to fuse strong text generation with deep visual understanding across images and video. It uses a mixture-of-experts architecture with roughly 22B active parameters drawn from a 235B total pool, a design that keeps inference efficient while preserving the reasoning capacity expected of a top-tier model. The release targets instruction-following workflows that demand long context, robust perception, and spatial grounding, positioning it as a flexible foundation for production multimodal systems.

In third-party evaluations, the model shows clear strengths in identification and document-style tasks while revealing softer spots in pure reasoning. On the Roboflow Vision Evals suite it averages about 65.9% across six tasks, climbing to roughly 90.6% on identification, 88.1% on OCR, and 87.6% on structured data extraction, with the weakest showing on visual reasoning. The n8n benchmark paints a complementary picture, awarding it a second-place overall score of 86, top placement on logic, and strong marks for cost efficiency and hallucination resistance. Practically, these profiles make Qwen3 VL 235B A22B Instruct a good fit for multilingual OCR, chart and table extraction, visual question answering, GUI automation, and agentic tool use, especially for teams that want an open-weight alternative to closed frontier vision-language systems.

Kilo Gatewayqwen/qwen3-vl-235b-a22b-instructqwen

Quick Info

Powered by
Provider
Kilo Gateway
Model key
qwen/qwen3-vl-235b-a22b-instruct
Release date
Sep 23, 2025
Last updated
Sep 23, 2025
Knowledge cutoff
2025-03-31
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.26
Output token cost
$1.04

Limits

Output tokens
32,768 tokens
Context window
131,072 tokens

Transparent token rates

Compare Qwen3 VL 235B A22B Instruct pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3 VL 235B A22B Instruct

OrcaRouter

Coverage

Qwen3-VL-235B-A22B-Instruct is an open-weight vision-language model from the Qwen team, described as the most capable VL model in the Qwen series. The weight repository is hosted on Hugging Face under Apache 2.0, and the model supports both Dense and MoE architectures with Instruct and Thinking editions for flexible deployment. This entry is the weight repository README for the exact 235B-A22B-Instruct variant. The model ships with a 256K native context extendable to 1M, 32-language OCR (up from 19), and enhanced spatial perception including 3D grounding for embodied AI. Architecture updates include Interleaved-MRoPE for video temporal reasoning, DeepStack ViT feature fusion, and text-timestamp alignment beyond T-RoPE. Capabilities include Visual Agent GUI operation, visual coding from images, and STEM/math reasoning with evidence-based answers.

Eden AI

CoverageBenchmark

Roboflow's Playground page for Qwen3 VL 235B A22B Instruct, published with pricing data updated September 18, 2026 and vision-eval scores updated September 5, 2026, confirms the model is a flagship multimodal vision-language model from Qwen/Alibaba that interleaves text and image inputs, supports very long contexts up On Roboflow's ground-truth Vision Evals benchmark, Qwen3 VL 235B A22B Instruct achieves an overall score of 65.9%, placing it 35th of 53 evaluated models, with an average cost per sample of $0.0007 (rank 11) and an average speed of 9.17 seconds per sample (rank 31), producing an average of 1.9K tokens per sample. The m

Videos about Qwen3 VL 235B A22B Instruct

More models around Qwen3 VL 235B A22B Instruct