Kilo Gateway
Qwen3 VL 32B Instruct pricing: $0.10/M input, $0.42/M output. Compare with 10 similar models, see benchmarks, and find the cheapest provider.
Model details
Qwen3 VL 32B Instruct is a decoder-only transformer model that integrates a specialized vision encoder to achieve deep fusion between text and visual data. Built with 33 billion parameters, the architecture utilizes Interleaved-MRoPE and DeepStack technologies to enhance spatial perception, allowing it to perform fine-grained object positioning and 3D grounding. This design intent focuses on creating a versatile agent capable of operating complex graphical user interfaces, generating code from visual inputs, and parsing long-form documents or hours-long video content with native support for extensive context windows.
The model benefits from a broad pretraining regimen that enables it to recognize a vast array of subjects, from landmarks and products to rare characters across 32 languages. Its lineage emphasizes robust performance in challenging conditions, such as low-light or blurred imagery, while maintaining strong logical reasoning for STEM and mathematical tasks. Designed for flexibility, the model supports customization through techniques like LoRA, making it a practical choice for developers looking to deploy specialized visual agents or high-accuracy OCR solutions in real-world environments.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Kilo Gateway
Qwen3 VL 32B Instruct pricing: $0.10/M input, $0.42/M output. Compare with 10 similar models, see benchmarks, and find the cheapest provider.
Kilo Gateway
When loading Qwen/Qwen3-VL-32B-Instruct with vLLM, for example when using TRL’s GRPOTrainer with use_vllm=True, an error of the form AttributeError: 'Qwen3VLTextConfig' object has no attribute 'tie...