Sulat.com
AI models
NanoGPT logo

Model details

Qwen3.5 0.8B

Qwen3.5 0.8B sits at the smallest end of the Qwen 3.5 family, a line of open-source multimodal models that combine language and vision understanding with tool-use and chain-of-thought capabilities. The Ollama distribution reports roughly 873 million parameters using a qwen35 architecture, offered as a Q8_0 quantized GGUF build of about 1.0 GB under an Apache License 2.0. On Qualcomm's optimized track, the same checkpoint appears as "unsloth/Qwen3.5-0.8B-GGUF" running through the GenieX llama.cpp runtime, with a more aggressive Q4_0 quantization that brings the footprint down to about 507 MB for on-device inference on Snapdragon-class hardware, supporting input combinations of image and text and producing text output.

In practice, the model targets lightweight assistants, content generation, and multimodal dialogue on edge devices such as phones, tablets, and IoT boards, rather than large-scale server workloads. The Qwen 3.5 family emphasizes architectural efficiency and multimodal learning, with the 0.8B variant intended to retain the family's vision, tool-calling, and reasoning features at a size suitable for constrained environments. Benchmark data from the Qualcomm distribution shows roughly 1,243 tokens per second prefilling and about 30.1 tokens per second decoding on the Snapdragon X2 Elite reference design at a 4,096-token NPU context window, a configuration that makes it appealing for developers who need a private, low-latency assistant or vision-aware helper running locally.

NanoGPTqwen3.5-0.8bqwen3.5

Quick Info

Powered by
Provider
NanoGPT
Model key
qwen3.5-0.8b
Release date
Aug 16, 2026
Last updated
Aug 16, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.06
Output token cost
$0.12

Limits

Input tokens
262,144 tokens
Output tokens
32,768 tokens
Context window
262,144 tokens

Latest news about Qwen3.5 0.8B

Videos about Qwen3.5 0.8B

Recent tweets and retweets from NanoGPT

More models around Qwen3.5 0.8B