Currently listed through these providers:
Model details
Qwen/Qwen3.5-9B
Qwen3.5-9B is positioned as a compact multimodal reasoning model that brings native tool calling and explicit chain-of-thought behavior into production workflows. The architecture pairs a hybrid Gated DeltaNet and Gated Attention design aimed at efficient inference with lower latency, while training combines early fusion over multimodal tokens with multi-token prediction and reinforcement learning across million-agent environments to keep the 9B parameter footprint competitive with larger peers. A thinking mode generates reasoning traces before answers, and native function calling targets agent reliability rather than just conversational quality.
Practically, the model is suited to long-running agents, document and video understanding, and global multilingual applications, with a 262K native context extendable beyond one million tokens via RoPE scaling and broad language coverage. Reported benchmark results highlight strong multimodal and agent skills, including OCRBench at 89.2%, VideoMME at 84.5%, MathVision at 78.9%, BFCL-V4 at 66.1%, TAU2-Bench at 79.1%, and MMMLU at 81.2%, illustrating balanced gains across vision, math, and tool orchestration. It fits teams that want a small-footprint base model that can be fine-tuned or deployed on demand for specialized pipelines without giving up agent-grade reasoning.
Quick Info
Powered by- Provider
- SiliconFlow
- Model key
- Qwen/Qwen3.5-9B
- Release date
- Mar 3, 2026
- Last updated
- Apr 24, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.10
- Output token cost
- $0.15
Limits
- Output tokens
- 262,144 tokens
- Context window
- 262,144 tokens