Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
SiliconFlow (China) logo

Model details

Qwen/Qwen3.5-4B

Qwen3.5-4B is part of the Qwen3.5 generation of compact foundation models released in early 2026, positioned as a small but multimodal-friendly option in the family. It is distributed as an open-weights checkpoint in Hugging Face Transformers format and is compatible with popular runtimes such as vLLM, SGLang, and KTransformers. The model card describes Qwen3.5 as a unified vision-language foundation trained with early fusion over multimodal tokens, aiming to match the reasoning and coding strengths of the text-only Qwen3 line while also surpassing earlier Qwen3-VL variants on visual understanding, agent tasks, and coding benchmarks.

Under the hood, the model uses a hybrid attention design that blends Gated Delta Networks with sparse Mixture-of-Experts layers, paired with a roughly 3:1 ratio of linear attention to full softmax attention for an efficient throughput-versus-quality balance. A native context window of about 262K tokens supports long documents and extended conversations, and the 4.66B-parameter footprint makes it well suited to on-device or cost-sensitive deployments where larger Qwen3.5 variants would be too heavy. Practically, this combination makes the model a practical pick for developers who want multimodal input handling plus reasoning and tool-use behavior in a lightweight open package.

SiliconFlow (China)Qwen/Qwen3.5-4Bqwen

Quick Info

Powered by
Provider
SiliconFlow (China)
Model key
Qwen/Qwen3.5-4B
Release date
Mar 3, 2026
Last updated
Mar 3, 2026
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
65,536 tokens
Context window
262,144 tokens

Latest news about Qwen/Qwen3.5-4B

SiliconFlow (China)

Coverage

The news blog specialized in Japanese culture, odd news, gadgets and all other funny stuffs. Updated everyday.

SiliconFlow (China)

CoverageBenchmark

Qwen3.5-4B is a 4 billion parameter vision-language model using Gated DeltaNet hybrid architecture with a 3:1 ratio of linear attention to full softmax attention. It supports 262K native context length and delivers strong performance for its size across knowledge, reasoning, coding, and multilingual tasks.

Videos about Qwen/Qwen3.5-4B

More models around Qwen/Qwen3.5-4B