Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Tempr Gateway logo

Model details

Qwen3.8 27B

Qwen3.8-27B is the dense, deployment-friendly entry in Alibaba's Qwen3.8 generation, a native vision-language model that ingests text and images and produces text outputs. It is published by the Qwen team on Hugging Face as a set of post-trained weights and configuration files in the Transformers layout and is engineered for compatibility with popular open-source serving stacks including Transformers, vLLM, SGLang, and TokenSpeed. The 27-billion-parameter design makes it tractable for local and self-hosted inference while still carrying the architectural lineage of Qwen3.5 that underpins the broader Qwen3.8 release.

Within the Qwen3.8 lineup, the 27B variant is positioned to consolidate gains seen across the Qwen3.5 and Qwen3.6 series into a single model that handles coding, professional work, research, and long-horizon agentic tasks, with the model card highlighting stronger autonomous planning and more reliable handling of environment feedback for end-to-end task completion. Flexible thinking control is part of the design intent, allowing developers to tune how much deliberate reasoning the model applies per request. A first-party hosted service on Qwen Cloud is planned to bring additional production features such as a 1M-token default context and built-in tools, broadening the model's fit for teams that want managed inference alongside the open-weights option.

Tempr Gatewaycerebras/qwen-3.8-27bqwen

Quick Info

Powered by
Provider
Tempr Gateway
Model key
cerebras/qwen-3.8-27b
Release date
Aug 14, 2026
Last updated
Sep 3, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.99
Output token cost
$1.49

Limits

Output tokens
40,960 tokens
Context window
131,072 tokens

Transparent token rates

Compare Qwen3.8 27B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.8 27B

Tempr

CoverageBenchmark

BigGo Finance reports that Qwen3.8-27B, released under Apache 2.0 by Alibaba's Tongyi Qianwen team, passed 1 million Hugging Face downloads within two days and inspired roughly 500 community quantized builds. The 27B dense model uses a hybrid Gated DeltaNet and Gated Attention architecture with a 262K native context window and multi-token prediction (MTP), and runs on consumer GPUs and workstations once quantized. The piece says NVIDIA, AMD, Cerebras, vLLM, and SGLang completed integrations shortly after release, while overseas developers pushed MTP speculative decoding, reasoning-effort tuning, quantization, and Apple Silicon work that yielded over 60% decoding-speed gains on some hardware. It situates Qwen3.8-27B inside broader Tongyi Qianwen claims of surpassing Qwen3.7-Plus and outperforming Claude Opus 4.6 Max on coding, agentic, and multimodal benchmarks, though those results trace to vendor-supplied numbers.

Tempr

CoverageBenchmark

Alibaba released Qwen3.8-27B with claimed Opus 4.6-level performance optimized for laptop inference. The open-weight release targets local deployment scenarios where developers run frontier-class models without cloud dependence. The piece frames Qwen3.8-27B as a local-inference milestone, positioning it against larger cloud-hosted competitors. It highlights practical implications for developers who want strong model quality on consumer hardware rather than relying on API-only access.

Tempr

CoverageBenchmark

Alibaba's Qwen team released Qwen3.8-27B on August 14, 2026, shipping as a 27.78B-parameter dense checkpoint under Apache 2.0 with a native 262,144-token context window and native text, image, and video inputs. Kingy.ai's launch-day review inspected the official model card, BF16 and FP8 artifacts, and GGUF inventory, and flagged that headlining benchmark gains (Terminal-Bench 2.1 to 73.0, DeepSWE 1.1 to 42.2, OSWorld-Verified to 84.3) come from Qwen itself and that the 1M-token claims require YaRN scaling. The review credits Qwen with large improvements over Qwen3.6-27B in agentic coding, computer use, and vision-language tasks without increasing decoder size, and rates it a strong candidate among locally deployable ~30B multimodal models. It also notes that YaRN-based long-context scaling can hurt shorter prompts, and that no independent benchmark reproduction was available on launch day, though Kingy later ran the exact model on an RTX 4090 24GB pairing for follow-up evidence.

Tempr

Coverage

A Hacker News thread confirms that Alibaba announced Qwen3.8-27B as an upcoming open-weight release alongside the Qwen3.8-Max launch. Community members noted Qwen3.6-27B and 35B as prior favorites for local use and expressed optimism that the 3.8 update would improve upon that lineage. Commenters discussed running earlier Qwen3.6 variants locally on 32GB machines for bulk non-code tasks, with one organization using Qwen-3.6-35B-A3B as an entry point for agent-first coding on RTX 5090s and AMD hardware. The thread provides grassroots developer perspective on Qwen local-model adoption.

Tempr

CoverageAnalysis

Alibaba's Tongyi Lab released Qwen3.8-27B on August 14, 2026, as a 27.78-billion-parameter dense multimodal language model under the Apache 2.0 license, making it free for runtime use. The release ships on Hugging Face and targets local deployment with a 24GB VRAM minimum requirement. The model uses a hybrid attention architecture built on Qwen3.5 with a 3:1 ratio of Gated DeltaNet linear attention to full attention across 64 transformer blocks, plus multi-token prediction for speculative decoding. It supports a 256K native context (extendable to 1M via YaRN), a 248,320-token multilingual vocabulary, and native multimodal input including text, image, and video.

Tempr

CoverageAnalysis

Qwen 3.8 27B was published on Hugging Face on August 14, 2026 under Apache 2.0 as a downloadable, self-hostable native vision-language dense model with 262,144 tokens of native context extendable to 1M via YaRN and a reasoning-effort dial defaulting to xhigh. Kie's deep-dive reports it fits on roughly 16–17GB of RAM/VRAM via Unsloth Dynamic GGUFs and hit ~10,000 downloads within 38 minutes and roughly 1 million downloads within 24 hours. The piece confirms the open weights, native multimodal architecture, and community excitement around the size-to-capability ratio, while noting that headline "beats Opus 4.6" coding claims trace to Qwen's own benchmark card with no independent verification in the signal set. It also flags that the xhigh default reasoning mode can dramatically over-think simple prompts, and identifies Qwen 3.8 27B as the open-weight sibling to the API-only Qwen 3.8-Max flagship rather than a frontier replacement.

Videos about Qwen3.8 27B

More models around Qwen3.8 27B