Qwen3.8 27B is a 27-billion-parameter dense model published by Alibaba's Qwen team as the open-weight companion to the much larger Qwen3.8 flagship, designed specifically to run on a single consumer or prosumer GPU rather than a cluster. Unlike the flagship, which ships under a bespoke licence, this 27B release is distributed under Apache 2.0, removing commercial-use restrictions and making it attractive for product teams, startups, and home-lab builders who want a permissive base for fine-tuning and redistribution. Its lineage carries forward the single-GPU workhorse role established by the predecessor Qwen3.6-27B, but adds native multimodality, accepting image and video inputs alongside text, so it can handle document understanding, STEM diagrams, and video analysis out of the box.
In practical terms, Qwen3.8 27B targets a sweet spot between capability and hardware footprint. At 4-bit quantisation it fits within roughly 17GB of VRAM, placing a 24GB card like the RTX 4090 squarely in play, while FP8 sits around 27GB for production single-card serving on an L40S or RTX 5090, and BF16 requires about 54GB on H100-class hardware for full-fidelity evaluation. Its native 262,144-token context window, extensible toward a million tokens, makes it well suited to long-document analysis and code repositories. Vendor-reported model-card benchmarks cited by independent coverage include GPQA Diamond at 89.2, LiveCodeBench v6 at 90.3, and SWE-bench Pro at 61.7, positioning it as a strong local all-rounder for coding, research, and agentic workflows without the cluster economics of larger open-weight competitors.