Amazon Bedrock
Qwen3 VL 235B A22B Instruct pricing: $0.20/M input, $0.88/M output. Compare with 10 similar models, see benchmarks, and find the cheapest provider.
Model details
Qwen3-VL-235B-A22B-Instruct is a flagship multimodal model designed to unify high-level text generation with sophisticated visual perception. Built on a mixture-of-experts architecture that utilizes approximately 22 billion active parameters out of a 235 billion total, the model is engineered to handle complex instruction-following across both text and visual inputs. Its design intent focuses on deep reasoning and spatial awareness, enabling it to perform tasks such as 2D and 3D grounding, precise object positioning, and detailed document parsing. By integrating these capabilities, the model serves as a versatile tool for applications ranging from automated GUI interaction to advanced multimodal research.
The model benefits from a broad, high-quality pretraining process that allows it to recognize a vast array of real-world categories, including landmarks, products, and diverse flora and fauna. Its training lineage emphasizes robust performance in multilingual OCR and visual coding, where it can translate sketches or mockups into functional code. With native support for long-context windows and second-level indexing for hours-long video, the model is well-suited for production environments requiring deep analysis of long documents or temporal video queries. These strengths, combined with its ability to act as a visual agent for PC and mobile interfaces, position it as a strong candidate for developers building embodied AI and complex, evidence-based reasoning systems.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Amazon Bedrock
Qwen3 VL 235B A22B Instruct pricing: $0.20/M input, $0.88/M output. Compare with 10 similar models, see benchmarks, and find the cheapest provider.