Currently listed through these providers:
Model details
Qwen3.5 0.8B
Qwen3.5 0.8B sits at the smallest end of the Qwen 3.5 family, a line of open-source multimodal models that combine language and vision understanding with tool-use and chain-of-thought capabilities. The Ollama distribution reports roughly 873 million parameters using a qwen35 architecture, offered as a Q8_0 quantized GGUF build of about 1.0 GB under an Apache License 2.0. On Qualcomm's optimized track, the same checkpoint appears as "unsloth/Qwen3.5-0.8B-GGUF" running through the GenieX llama.cpp runtime, with a more aggressive Q4_0 quantization that brings the footprint down to about 507 MB for on-device inference on Snapdragon-class hardware, supporting input combinations of image and text and producing text output.
In practice, the model targets lightweight assistants, content generation, and multimodal dialogue on edge devices such as phones, tablets, and IoT boards, rather than large-scale server workloads. The Qwen 3.5 family emphasizes architectural efficiency and multimodal learning, with the 0.8B variant intended to retain the family's vision, tool-calling, and reasoning features at a size suitable for constrained environments. Benchmark data from the Qualcomm distribution shows roughly 1,243 tokens per second prefilling and about 30.1 tokens per second decoding on the Snapdragon X2 Elite reference design at a 4,096-token NPU context window, a configuration that makes it appealing for developers who need a private, low-latency assistant or vision-aware helper running locally.
Quick Info
Powered by- Provider
- NanoGPT
- Model key
- qwen3.5-0.8b
- Release date
- Aug 16, 2026
- Last updated
- Aug 16, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.06
- Output token cost
- $0.12
Limits
- Input tokens
- 262,144 tokens
- Output tokens
- 32,768 tokens
- Context window
- 262,144 tokens