Sulat.com
AI models
LLM Tech logo

Model details

Qwen3.8 27B

Qwen3.8-27B is the dense 27-billion-parameter member of the Qwen3.8 generation, designed to bring flagship-level capability into a more deployment-friendly size. It is built on the Qwen3.5 architectural foundation and pairs a causal language model with a separate vision encoder, making it a native vision-language system that understands text, images, and video. The model is positioned for coding, professional work, research, and long-horizon agentic tasks, and ships with stronger autonomous planning and more reliable end-to-end task completion compared to earlier Qwen3.5 and Qwen3.6 releases, along with broader harness and tooling compatibility.

At the architecture level, Qwen3.8-27B uses a hybrid attention backbone of 64 layers: 16 layers run full gated attention on a fixed interval while the remaining 48 layers run linear attention via Gated DeltaNet with a constant recurrent state. It also includes a built-in Multi-Token Prediction (MTP) draft head and a native 262,144-token context window that can be extended toward one million tokens through RoPE scaling. Thinking mode is enabled by default, with tunable reasoning effort and preserved reasoning context across multi-step work, giving practitioners flexible control over depth, latency, and cost in real workloads.

LLM Techunsloth/Qwen3.8-27B-NVFP4qwen

Quick Info

Powered by
Provider
LLM Tech
Model key
unsloth/Qwen3.8-27B-NVFP4
Release date
Aug 14, 2026
Last updated
Aug 14, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.25
Output token cost
$2.09

Limits

Output tokens
32,768 tokens
Context window
262,144 tokens

Latest news about Qwen3.8 27B

Videos about Qwen3.8 27B

More models around Qwen3.8 27B