Currently listed through these providers:
Model details
Qwen3.8 27B
Qwen3.8-27B is the dense 27-billion-parameter member of the Qwen3.8 generation, designed to bring flagship-level capability into a more deployment-friendly size. It is built on the Qwen3.5 architectural foundation and pairs a causal language model with a separate vision encoder, making it a native vision-language system that understands text, images, and video. The model is positioned for coding, professional work, research, and long-horizon agentic tasks, and ships with stronger autonomous planning and more reliable end-to-end task completion compared to earlier Qwen3.5 and Qwen3.6 releases, along with broader harness and tooling compatibility.
At the architecture level, Qwen3.8-27B uses a hybrid attention backbone of 64 layers: 16 layers run full gated attention on a fixed interval while the remaining 48 layers run linear attention via Gated DeltaNet with a constant recurrent state. It also includes a built-in Multi-Token Prediction (MTP) draft head and a native 262,144-token context window that can be extended toward one million tokens through RoPE scaling. Thinking mode is enabled by default, with tunable reasoning effort and preserved reasoning context across multi-step work, giving practitioners flexible control over depth, latency, and cost in real workloads.
Quick Info
Powered by- Provider
- LLM Tech
- Model key
- unsloth/Qwen3.8-27B-NVFP4
- Release date
- Aug 14, 2026
- Last updated
- Aug 14, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.25
- Output token cost
- $2.09
Limits
- Output tokens
- 32,768 tokens
- Context window
- 262,144 tokens
Latest news about Qwen3.8 27B
Videos about Qwen3.8 27B
More models around Qwen3.8 27B
This exact model name is also listed by 27 other providers.