OrcaRouter
Qwen3-VL-235B-A22B-Instruct is an open-weight vision-language model from the Qwen team, described as the most capable VL model in the Qwen series. The weight repository is hosted on Hugging Face under Apache 2.0, and the model supports both Dense and MoE architectures with Instruct and Thinking editions for flexible deployment. This entry is the weight repository README for the exact 235B-A22B-Instruct variant. The model ships with a 256K native context extendable to 1M, 32-language OCR (up from 19), and enhanced spatial perception including 3D grounding for embodied AI. Architecture updates include Interleaved-MRoPE for video temporal reasoning, DeepStack ViT feature fusion, and text-timestamp alignment beyond T-RoPE. Capabilities include Visual Agent GUI operation, visual coding from images, and STEM/math reasoning with evidence-based answers.