Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
routing.run logo

Model details

Qwen3.5 9B

Qwen3.5 9B is an open-weight release from the Qwen family, distributed as a post-trained model with weights and configuration files in the Hugging Face Transformers format and compatibility with inference stacks such as vLLM, SGLang, and KTransformers. The creator describes Qwen3.5 as a unified vision-language foundation that uses early-fusion training over multimodal tokens, claiming cross-generational parity with Qwen3 and stronger results than Qwen3-VL on reasoning, coding, agents, and visual understanding benchmarks. The same release notes frame the family as a step forward in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility, with expanded coverage of 201 languages and dialects.

Under the hood, Qwen3.5 pairs Gated Delta Networks with a sparse Mixture-of-Experts design to deliver high-throughput inference with minimal latency overhead, and the creator points to near-100% multimodal training efficiency relative to text-only runs alongside asynchronous reinforcement learning infrastructure for agent training at scale. These traits position the 9B variant as a practical mid-size option for developers who want a single open-weight model that can handle visual understanding alongside text reasoning and tool use, rather than stitching together separate vision and language models. Deployment is straightforward across the usual open-source runtimes, and the same model is also packaged for ROCm-based AMD Inference Microservice serving on Instinct, Radeon Pro, and EPYC hardware with an OpenAI-compatible API, broadening the environments where it can be put to work.

routing.runqwen3.5-9bqwen

Quick Info

Powered by
Provider
routing.run
Model key
qwen3.5-9b
Release date
Feb 23, 2026
Last updated
Feb 23, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.16
Output token cost
$0.48

Limits

Output tokens
32,000 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen3.5 9B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.5 9B

No articles yet. Fetch the latest news to show it here.

Videos about Qwen3.5 9B

More models around Qwen3.5 9B