Mixlayer
I ran benchmarks of Qwen3.5 35b through 0.8b models using all different quants and compared performance with flash attention enabled vs disabled. Using llama.cpp and Unsloth’s UD_K_XL versions. Specifically for the Qwen…
Model details
Qwen3.5 9B is a dense, 9-billion-parameter multimodal model built to combine reasoning, coding, and visual understanding in a relatively compact architecture. Its unified vision-language design uses early fusion, so visual tokens are incorporated alongside text rather than handled as a separate pipeline. This makes the model especially relevant for applications that must reason over documents, diagrams, screenshots, or other visual material while retaining language-based problem solving.
The model is designed for tool-oriented and reasoning-enabled workflows, with a native context of 262,144 tokens that supports substantial conversational or project context. Its long-context design and visual-language integration suit coding assistants, document analysis, visual question answering, and agent-like tasks that combine images with instructions. Benchmarks reported for the model show strong graduate-level science reasoning, but more mixed results on long-context and agentic evaluations, so it is best viewed as a balanced compact multimodal option rather than a specialist in every reasoning task.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Mixlayer
I ran benchmarks of Qwen3.5 35b through 0.8b models using all different quants and compared performance with flash attention enabled vs disabled. Using llama.cpp and Unsloth’s UD_K_XL versions. Specifically for the Qwen…
TensorX
The BenchmarkList page for Qwen3.5-9B lists 101 benchmark rows and prices the model at $0.10 per 1M input tokens and $0.15 per 1M output tokens, situating it alongside comparison rows for Qwen Fable 5.1, Claude Fable 5.1, Claude Opus 5, Kimi K3, Qwen3.8-Flash-Next, Qwen3.8-2.4T-A95B, GLM 5.3, and GLM 5.3 Flash. In the The page distinguishes Qwen3.5-9B from sibling and successor variants by tagging the exact slug qwen-qwen3.5-9b and listing it next to Qwen3.8-Flash-Next and Qwen3.8-2.4T-A95B, allowing direct comparison against newer Qwen3.8 releases on the same benchmark rows. It functions as a model-focused benchmark hub rather than
Mixlayer
OpenRouter's first-party catalog page for qwen/qwen3.5-9b documents the model as a multimodal foundation model in the Qwen 3.5 family, using a unified vision-language design with early fusion of multimodal tokens so it can reason across text and images in a single context. The page confirms a 262,144-token context wind The listing also exposes routing modes (Balanced, Nitro, Exacto), a "Compare" playground, and an API key flow, giving developers concrete integration context for using qwen/qwen3.5-9B behind an OpenAI-compatible interface. Capabilities called out include reasoning, tool calling, structured output, attachments, open wei