Currently listed through these providers:
Model details
Qwen3-VL 235B
Qwen3-VL 235B sits at the top of the Qwen vision-language family as a mixture-of-experts design that activates 22 billion of its 235 billion total parameters per inference, combining the efficiency of sparse routing with the capacity needed for demanding multimodal workloads. The served STACKIT variant is an 8-bit quantized build of the base instruct model, preserving the open-weight lineage while reducing the GPU footprint relative to the full-precision release. Independent providers position this configuration as their flagship vision-language offering, and the underlying instruct checkpoint was rated as the top open model for text understanding on lmarena.ai according to SGLang documentation, signaling that the model does not sacrifice language quality for its visual strengths.
In practical terms, the model is built for agentic and visually grounded applications: GUI interaction on PCs and mobile devices, visual coding that turns images or video into Draw.io diagrams and front-end code, spatial perception with 2D and emerging 3D grounding for embodied tasks, and OCR that spans 32 languages including challenging document conditions. A context window that stretches into the hundreds of thousands of tokens lets it ingest long videos, lengthy PDFs, and multi-document reasoning sessions in a single pass. The combination of MoE efficiency, expanded visual recognition from broader pretraining, and tool-use support makes it a strong fit for teams building autonomous assistants, document intelligence pipelines, and multimodal research prototypes that need both deep perception and reliable instruction following.
Quick Info
Powered by- Provider
- STACKIT
- Model key
- Qwen/Qwen3-VL-235B-A22B-Instruct-FP8
- Release date
- Nov 1, 2024
- Last updated
- Nov 1, 2024
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $1.76
- Output token cost
- $2.05
Limits
- Output tokens
- 16,384 tokens
- Context window
- 218,000 tokens