Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Tempr Gateway logo

Model details

Qwen3.8 27B

Qwen3.8-27B is a 27-billion-parameter dense model released by Alibaba's Qwen team as part of the Qwen3.8 generation, the most capable open-model family release to date according to the official model card. It is built on the architectural foundation of Qwen3.5 and is positioned as a compact, deployment-friendly alternative within the family, bringing the generation's advances in coding, professional work, research, and long-horizon agentic work to a single dense checkpoint that developers can download and fine-tune on their own hardware.

Designed as a native vision-language model, Qwen3.8-27B understands images and supports flexible thinking control, which lets users steer how much deliberate reasoning the model applies to a given prompt. The model ships with open weights compatible with popular inference stacks including Hugging Face Transformers, vLLM, SGLang, and TokenSpeed, and the Qwen team highlights stronger autonomous planning and better handling of environment feedback for more reliable end-to-end task completion. Practically, it fits developers and teams who want a self-hosted model that can carry multi-step coding and agentic workflows, with an upcoming hosted version on Qwen Cloud offering expanded context and built-in tools for production use.

Tempr Gatewaygroq/qwen/qwen3.8-27bqwen

Quick Info

Powered by
Provider
Tempr Gateway
Model key
groq/qwen/qwen3.8-27b
Release date
Aug 14, 2026
Last updated
Aug 14, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.80
Output token cost
$4.00

Limits

Output tokens
16,384 tokens
Context window
131,042 tokens

Transparent token rates

Compare Qwen3.8 27B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.8 27B

Tempr

CoverageBenchmark

Regolo.ai provides a head-to-head benchmark comparison of Qwen3.8-27B against Claude Opus 4.6 Max across 24 benchmarks from the official Alibaba model card. Qwen3.8-27B wins 16, with largest gains in agentic coding (SWE-bench Pro 61.7 vs 53.4), instruction following (IFBench 79.5 vs 62.5), and computer use (OSWorld-Verified 84.3 vs 72.7). It loses on knowledge-heavy reasoning such as Humanity's Last Exam (30.8 vs 40.0) and GPQA Diamond. The article confirms Qwen3.8-27B is a dense 27-billion-parameter open-weight vision-language model with 28 billion parameters including the vision encoder, running on a single 24GB GPU at 4-bit quantization. It also notes that against Muse Glimmer-30B, Qwen3.8-27B wins every overlapping benchmark row. All figures are explicitly vendor-reported by Alibaba as of August 14, 2026, with no independent reproduction existing at launch.

Tempr

CoverageBenchmark

Qwen3.8 27B is an open-weight dense model from the Qwen team highlighted in The New Stack as delivering near Opus 4.6-class performance while running on a laptop. The piece frames the release as a local-inference milestone, giving developers a strong coding-capable model without cloud dependence. This positions 27B as a practical desktop alternative to larger hosted systems. The article confirms Qwen3.8 27B targets local deployment and professional workflows, reinforcing its open-weight availability for the developer community. It underlines Alibaba's Qwen team's continued emphasis on efficiency at the 27B parameter scale. This makes the model relevant for teams prioritizing on-device inference and cost control.

Tempr

CoverageBenchmark

Qwen3.8-27B is a dense 27.78-billion-parameter multimodal model released by Alibaba's Qwen team on August 14, 2026, under Apache 2.0. The Kingy.ai launch-day review confirms a native 262,144-token context window, with text, image, and video input support. Vendor-reported benchmark deltas versus Qwen3.6-27B include Terminal-Bench 2.1 rising from 63.4 to 73.0, DeepSWE 1.1 from 13.3 to 42.2, and OSWorld-Verified from 63.9 to 84.3. The review flags that all launch scores are self-reported by Qwen and that several benchmarks are in-house or modified, with the SWE-bench Pro Opus comparison imported rather than rerun. It also warns that local-hardware claims often ignore KV cache memory and that the 1M-token context requires YaRN scaling rather than the native window. Official BF16 and FP8 artifacts plus third-party GGUF files were inspected, but no inference was run at launch due to the Qwen Cloud endpoint being marked "coming soon."

Tempr

CoverageBenchmark

OpenRouter's listing confirms Qwen3.8 27B is an open-weight dense vision-language model from Qwen, suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks with toggleable thinking. It supports a 262K-token context and was released on August 14, 2026. These attributes establish the model's core capability profile. The page surfaces provider-by-provider routing metrics including latency (0.41s P50), throughput (34 tok/s), and benchmark percentages on GPQA Diamond, TAU-Bench, and Darkbloom across hosts such as ModelRun, Cloudflare, DekaLLM, and Parasail. While useful for deployment decisions, these figures describe serving-provider performance rather than intrinsic model characteristics.

Tempr

CoverageAnalysis

Local AI Zone provides a comprehensive technical analysis of Qwen3.8-27B, confirming it as a 27.78-billion-parameter dense multimodal model from Alibaba's Tongyi Lab released August 14, 2026 under Apache 2.0. The model uses a hybrid attention architecture with 48 Gated DeltaNet linear-attention layers and 16 full-attention layers, totaling 64 transformer blocks with a 5,120 hidden dimension. Native context is 262,144 tokens, extendable to 1 million via YaRN. The analysis reports Qwen3.8-27B achieves competitive performance against models 10-15x its size, running on as little as 24GB VRAM. It outperforms Meta's Muse Glimmer (30B) across all 8 direct comparison benchmarks and surpasses Claude Opus 4.6 on 15 of 19 overlapping tests. Key architectural features include multi-token prediction for speculative decoding, 24 query heads with 4 KV heads, and a fused attention output gate, signaling a new efficiency frontier for local AI deployment.

Videos about Qwen3.8 27B

More models around Qwen3.8 27B