Model details
Qwen3.5 4B
Qwen3.5 4B sits inside a multimodal family built around a unified vision-language foundation, where early-fusion training on multimodal tokens lets the model reach parity with the prior Qwen3 generation and outperform the Qwen3-VL line across reasoning, coding, agent, and visual understanding benchmarks. The model card frames Qwen3.5 as a step forward in multimodal learning and global accessibility, signaling an intent to serve developers and enterprises that need a compact base for both text and image understanding rather than a single-modality specialist. This positioning makes the 4B variant a practical entry point into the family's capabilities without committing to the larger sizes.
Under the hood, Qwen3.5 relies on an efficient hybrid architecture that pairs Gated Delta Networks with a sparse Mixture-of-Experts design, aiming for high-throughput inference with minimal latency overhead. Training is reported to scale reinforcement learning across million-agent environments with progressively complex task distributions, and the infrastructure claims near-complete multimodal training efficiency relative to text-only runs. The Ollama distribution confirms a roughly 4.66B-parameter build released under Apache 2.0, compatible with runtimes like vLLM, SGLang, and KTransformers, which gives teams flexibility in deployment while keeping the weights openly available for fine-tuning and local experimentation.
Quick Info
Powered by- Provider
- EmpirioLabs AI
- Model key
- qwen3-5-4b
- Release date
- Mar 2, 2026
- Last updated
- Mar 2, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.04
- Output token cost
- $0.07
Limits
- Output tokens
- 32,768 tokens
- Context window
- 262,144 tokens
Latest news about Qwen3.5 4B
No articles yet. Fetch the latest news to show it here.