Sulat.com
AI models
SiliconFlow logo

Model details

Qwen/Qwen3.5-9B

Qwen3.5-9B is positioned as a compact multimodal reasoning model that brings native tool calling and explicit chain-of-thought behavior into production workflows. The architecture pairs a hybrid Gated DeltaNet and Gated Attention design aimed at efficient inference with lower latency, while training combines early fusion over multimodal tokens with multi-token prediction and reinforcement learning across million-agent environments to keep the 9B parameter footprint competitive with larger peers. A thinking mode generates reasoning traces before answers, and native function calling targets agent reliability rather than just conversational quality.

Practically, the model is suited to long-running agents, document and video understanding, and global multilingual applications, with a 262K native context extendable beyond one million tokens via RoPE scaling and broad language coverage. Reported benchmark results highlight strong multimodal and agent skills, including OCRBench at 89.2%, VideoMME at 84.5%, MathVision at 78.9%, BFCL-V4 at 66.1%, TAU2-Bench at 79.1%, and MMMLU at 81.2%, illustrating balanced gains across vision, math, and tool orchestration. It fits teams that want a small-footprint base model that can be fine-tuned or deployed on demand for specialized pipelines without giving up agent-grade reasoning.

SiliconFlowQwen/Qwen3.5-9Bqwen

Quick Info

Powered by
Provider
SiliconFlow
Model key
Qwen/Qwen3.5-9B
Release date
Mar 3, 2026
Last updated
Apr 24, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.10
Output token cost
$0.15

Limits

Output tokens
262,144 tokens
Context window
262,144 tokens

Latest news about Qwen/Qwen3.5-9B

Videos about Qwen/Qwen3.5-9B

Recent tweets and retweets from SiliconFlow

More models around Qwen/Qwen3.5-9B