SiliconFlow (China)
The news blog specialized in Japanese culture, odd news, gadgets and all other funny stuffs. Updated everyday.
Model details
Qwen3.5-4B is part of the Qwen3.5 generation of compact foundation models released in early 2026, positioned as a small but multimodal-friendly option in the family. It is distributed as an open-weights checkpoint in Hugging Face Transformers format and is compatible with popular runtimes such as vLLM, SGLang, and KTransformers. The model card describes Qwen3.5 as a unified vision-language foundation trained with early fusion over multimodal tokens, aiming to match the reasoning and coding strengths of the text-only Qwen3 line while also surpassing earlier Qwen3-VL variants on visual understanding, agent tasks, and coding benchmarks.
Under the hood, the model uses a hybrid attention design that blends Gated Delta Networks with sparse Mixture-of-Experts layers, paired with a roughly 3:1 ratio of linear attention to full softmax attention for an efficient throughput-versus-quality balance. A native context window of about 262K tokens supports long documents and extended conversations, and the 4.66B-parameter footprint makes it well suited to on-device or cost-sensitive deployments where larger Qwen3.5 variants would be too heavy. Practically, this combination makes the model a practical pick for developers who want multimodal input handling plus reasoning and tool-use behavior in a lightweight open package.
A provider subscription or plan supersedes token-based pricing for this model.
SiliconFlow (China)
The news blog specialized in Japanese culture, odd news, gadgets and all other funny stuffs. Updated everyday.
SiliconFlow (China)
Qwen3.5-4B is a 4 billion parameter vision-language model using Gated DeltaNet hybrid architecture with a 3:1 ratio of linear attention to full softmax attention. It supports 262K native context length and delivers strong performance for its size across knowledge, reasoning, coding, and multilingual tasks.