Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
SiliconFlow (China) logo

Model details

Qwen/Qwen3.5-9B

Qwen3.5-9B is positioned as a dense multimodal reasoning model within the broader Qwen3.5 family, bringing together text, image, and video understanding in a single 9 billion parameter package. According to the Together AI model card, the architecture combines a hybrid Gated DeltaNet and Gated Attention design intended for efficient inference, and the model keeps a native context window that can be extended well beyond its base length with RoPE scaling. Vision is handled by a dedicated encoder that feeds into the same backbone, and the description explicitly highlights early fusion training on multimodal tokens, large-scale reinforcement learning across million-agent environments, and multi-token prediction training as part of the recipe.

In practical terms, the model is aimed at agentic and long-context workloads rather than narrow single-task use. Together AI reports function-calling results of 66.1% on BFCL-V4 and 79.1% on TAU2-Bench for production agents, alongside multimodal reasoning scores such as 89.2% on OCRBench, 84.5% on VideoMME, and 78.9% on MathVision, plus broad language coverage at 81.2% on MMMLU. The Together card also highlights an explicit thinking mode that produces reasoning traces before final answers. Because it is open weights and small enough to run locally with modest memory, as reflected in the LM Studio listing, it is a reasonable fit for teams that want a self-hostable model capable of grounded reasoning over long documents, tool-driven workflows, and multimodal inputs without the cost of a flagship-scale system.

SiliconFlow (China)Qwen/Qwen3.5-9Bqwen

Quick Info

Powered by
Provider
SiliconFlow (China)
Model key
Qwen/Qwen3.5-9B
Release date
Mar 3, 2026
Last updated
Mar 3, 2026
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.22
Output token cost
$1.74

Limits

Output tokens
65,536 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen/Qwen3.5-9B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen/Qwen3.5-9B

SiliconFlow (China)

CoverageBenchmark

I ran benchmarks of Qwen3.5 35b through 0.8b models using all different quants and compared performance with flash attention enabled vs disabled. Using llama.cpp and Unsloth’s UD_K_XL versions. Specifically for the Qwen…

SiliconFlow (China)

CoverageBenchmark

The Together AI model page explicitly names Qwen/Qwen3.5-9B and documents its architecture: a hybrid Gated DeltaNet and Gated Attention design with 9 billion parameters, a 262,144-token native context window extensible to 1M+ tokens via RoPE scaling, and a vision encoder supporting text, image, and video inputs through The page lists a dense benchmark slate for the model: 89.2% on OCRBench, 84.5% on VideoMME, 78.9% on MathVision, 70.1% on MMMU-Pro, 66.1% on BFCL-V4, 79.1% on TAU2-Bench, 65.6% on LiveCodeBench v6, 81.2% on MMMLU across 201 languages, 63.0% on AA-LCR, and 55.2% on LongBench v2. API access is available via the Together

Videos about Qwen/Qwen3.5-9B

More models around Qwen/Qwen3.5-9B