Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
QVAC logo

Model details

Qwen3.5 0.8B

Qwen3.5 0.8B sits within the broader Qwen3.5 family, which Alibaba describes as integrating multimodal learning, architectural efficiency, reinforcement learning scale, and global linguistic coverage. The family-level announcement highlights a unified vision-language foundation built through early fusion training on multimodal tokens, an efficient hybrid architecture pairing Gated Delta Networks with sparse Mixture-of-Experts, and reinforcement learning scaled across million-agent environments with progressively complex task distributions. CanIRun.ai positions the 0.8B variant specifically for chat and edge use cases, making it the smallest, most deployment-friendly entry point into the Qwen3.5 lineup for developers who need local or on-device inference rather than a flagship-scale foundation model.

On the practical side, Qwen3.5 0.8B is available through community channels such as Ollama, where CanIRun.ai records more than 406,000 downloads and 317 likes, alongside multiple GGUF quantizations that keep the footprint small enough for consumer hardware. Q4_K_M lands near 0.9 GB of VRAM at “good” quality, while Q6_K and Q8_0 reach roughly 1.1 GB and 1.3 GB respectively at “excellent” quality, giving developers an easy trade-off between memory budget and response fidelity. A DeepInfra technical blog post is also published with the stated intent of measuring latency, throughput, and cost for the same checkpoint, although the supplied excerpt only contains page navigation rather than measurable results, so any concrete performance numbers should be verified against the live post before being quoted.

QVACqwen3.5-0.8bqwen

Quick Info

Powered by
Provider
QVAC
Model key
qwen3.5-0.8b
Release date
Nov 1, 2025
Last updated
Nov 1, 2025
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
8,192 tokens
Context window
32,768 tokens

Latest news about Qwen3.5 0.8B

QVAC

Coverage

Qwen3.5-0.8B is Alibaba Cloud's ultra-compact multimodal foundation model with 800 million parameters, released on 23 February 2026 under the Apache 2.0 license, according to the apxml.com model specification page. It uses a dense architecture with a hybrid attention pattern combining Gated Delta Networks and Gated Att The model features grouped-query attention (8Q/2KV heads, head dim 256), RoPE with theta 10,000,000, RMSNorm, SwiGLU activation, and a 75% linear-attention ratio, along with multi-token prediction training and unified vision-language capabilities. According to the same page, Qwen3.5-0.8B is intended for prototyping, fi

Videos about Qwen3.5 0.8B

More models around Qwen3.5 0.8B