Currently listed through these providers:
Model details
Qwen3.5 9B
Qwen3.5 9B is an open-weight release from the Qwen family, distributed as a post-trained model with weights and configuration files in the Hugging Face Transformers format and compatibility with inference stacks such as vLLM, SGLang, and KTransformers. The creator describes Qwen3.5 as a unified vision-language foundation that uses early-fusion training over multimodal tokens, claiming cross-generational parity with Qwen3 and stronger results than Qwen3-VL on reasoning, coding, agents, and visual understanding benchmarks. The same release notes frame the family as a step forward in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility, with expanded coverage of 201 languages and dialects.
Under the hood, Qwen3.5 pairs Gated Delta Networks with a sparse Mixture-of-Experts design to deliver high-throughput inference with minimal latency overhead, and the creator points to near-100% multimodal training efficiency relative to text-only runs alongside asynchronous reinforcement learning infrastructure for agent training at scale. These traits position the 9B variant as a practical mid-size option for developers who want a single open-weight model that can handle visual understanding alongside text reasoning and tool use, rather than stitching together separate vision and language models. Deployment is straightforward across the usual open-source runtimes, and the same model is also packaged for ROCm-based AMD Inference Microservice serving on Instinct, Radeon Pro, and EPYC hardware with an OpenAI-compatible API, broadening the environments where it can be put to work.
Quick Info
Powered by- Provider
- routing.run
- Model key
- qwen3.5-9b
- Release date
- Feb 23, 2026
- Last updated
- Feb 23, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.16
- Output token cost
- $0.48
Limits
- Output tokens
- 32,000 tokens
- Context window
- 262,144 tokens
Transparent token rates
Compare Qwen3.5 9B pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Qwen3.5 9B
No articles yet. Fetch the latest news to show it here.
