Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
NovitaAI logo

Model details

Qwen3 VL 235B A22B Instruct

Qwen3 VL 235B A22B Instruct is a flagship multimodal vision-language model released in September 2025 by the Qwen team at Alibaba Cloud, built to fuse strong text generation with deep visual understanding across images and video. It uses a mixture-of-experts architecture with roughly 22B active parameters drawn from a 235B total pool, a design that keeps inference efficient while preserving the reasoning capacity expected of a top-tier model. The release targets instruction-following workflows that demand long context, robust perception, and spatial grounding, positioning it as a flexible foundation for production multimodal systems.

In third-party evaluations, the model shows clear strengths in identification and document-style tasks while revealing softer spots in pure reasoning. On the Roboflow Vision Evals suite it averages about 65.9% across six tasks, climbing to roughly 90.6% on identification, 88.1% on OCR, and 87.6% on structured data extraction, with the weakest showing on visual reasoning. The n8n benchmark paints a complementary picture, awarding it a second-place overall score of 86, top placement on logic, and strong marks for cost efficiency and hallucination resistance. Practically, these profiles make Qwen3 VL 235B A22B Instruct a good fit for multilingual OCR, chart and table extraction, visual question answering, GUI automation, and agentic tool use, especially for teams that want an open-weight alternative to closed frontier vision-language systems.

NovitaAIqwen/qwen3-vl-235b-a22b-instruct

Quick Info

Powered by
Provider
NovitaAI
Model key
qwen/qwen3-vl-235b-a22b-instruct
Release date
Sep 24, 2025
Last updated
Sep 24, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.30
Output token cost
$1.50

Limits

Output tokens
32,768 tokens
Context window
131,072 tokens

Latest news about Qwen3 VL 235B A22B Instruct

OrcaRouter

Coverage

Qwen3-VL-235B-A22B-Instruct is an open-weight vision-language model from the Qwen team, described as the most capable VL model in the Qwen series. The weight repository is hosted on Hugging Face under Apache 2.0, and the model supports both Dense and MoE architectures with Instruct and Thinking editions for flexible deployment. This entry is the weight repository README for the exact 235B-A22B-Instruct variant. The model ships with a 256K native context extendable to 1M, 32-language OCR (up from 19), and enhanced spatial perception including 3D grounding for embodied AI. Architecture updates include Interleaved-MRoPE for video temporal reasoning, DeepStack ViT feature fusion, and text-timestamp alignment beyond T-RoPE. Capabilities include Visual Agent GUI operation, visual coding from images, and STEM/math reasoning with evidence-based answers.

Eden AI

CoverageBenchmark

Roboflow's Playground page for Qwen3 VL 235B A22B Instruct, published with pricing data updated September 18, 2026 and vision-eval scores updated September 5, 2026, confirms the model is a flagship multimodal vision-language model from Qwen/Alibaba that interleaves text and image inputs, supports very long contexts up On Roboflow's ground-truth Vision Evals benchmark, Qwen3 VL 235B A22B Instruct achieves an overall score of 65.9%, placing it 35th of 53 evaluated models, with an average cost per sample of $0.0007 (rank 11) and an average speed of 9.17 seconds per sample (rank 31), producing an average of 1.9K tokens per sample. The m

NovitaAI

CoverageBenchmark

n8n's AI Benchmark page for the exact subject model (Qwen3-VL-235B-A22B Instruct) describes it as an open-weight multimodal model unifying strong text generation with visual understanding across images and video, targeting VQA, document parsing, chart/table extraction, and multilingual OCR. It emphasizes robust percept n8n's composite benchmark scores rank the model 2nd overall with a score of 86 (behind xAI Grok 4 Fast), with standout placements in Logic (96, 1st), Hallucination (91, 4th), and Cost (98, 3rd). Pricing listed on the page appears as $0.0000002 per 1K prompt tokens and $0.0000012 per 1K completion tokens, which the rese

Videos about Qwen3 VL 235B A22B Instruct