Currently listed through these providers:
Model details
Qwen3.5 9B
Qwen3.5 9B sits inside Alibaba's broader Qwen3.5 release, which the team describes as a unified vision-language foundation trained with early fusion over multimodal tokens. The official model card states that this approach reaches cross-generational parity with the earlier Qwen3 line and outperforms Qwen3-VL variants on reasoning, coding, agent, and visual understanding benchmarks, framing the 9B model as a compact multimodal generalist rather than a text-only workhorse. Its multimodal scope pairs naturally with the model's reasoning and tool-calling capabilities, letting developers route text, image, and video inputs through a single checkpoint for tasks like document analysis, visual question answering, and tool-assisted workflows.
Under the hood, Qwen3.5 9B relies on a hybrid architecture that pairs Gated Delta Networks with a sparse Mixture-of-Experts design, which the Qwen team highlights as a path to high-throughput inference with minimal latency overhead. The weights are published openly on Hugging Face in a Transformers-compatible format that also runs under vLLM, SGLang, and KTransformers, and the model is mirrored on Ollama as a roughly 9.65B-parameter Q4_K_M build weighing about 6.6GB under Apache 2.0. That combination of an efficient hybrid backbone, open licensing, and broad runtime support makes the 9B variant a practical choice for teams that want a capable, locally deployable multimodal model with room to scale into larger Qwen3.5 siblings when more capacity is available.
Quick Info
Powered by- Provider
- OpenRouter
- Model key
- qwen/qwen3.5-9b
- Release date
- Feb 23, 2026
- Last updated
- Feb 23, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.10
- Output token cost
- $0.15
Limits
- Output tokens
- 235,929 tokens
- Context window
- 262,144 tokens