Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Cerebras logo

Model details

Qwen3.8 27B

Qwen3.8-27B is a 27-billion-parameter dense model from Alibaba's Qwen open-model family, released with publicly available weights on Hugging Face and positioned as a compact, deployment-friendly option within the Qwen3.8 generation. It is built on the architectural foundation of Qwen3.5, continuing the lineage that followed the widely adopted Qwen3.5 and Qwen3.6 series, and is described as the most capable generation in that open family so far.

As a native vision-language model, Qwen3.8-27B processes images alongside text and is tuned with flexible thinking control for multi-step problem solving. The Qwen3.8 series emphasizes substantial gains in coding, professional work, research, and long-horizon agentic tasks, with stronger autonomous planning and improved handling of environment feedback for more reliable end-to-end execution. The post-trained checkpoint ships as standard weights compatible with Hugging Face Transformers, vLLM, SGLang, and TokenSpeed, making it practical for local and self-hosted deployments that need a balance of capability and footprint.

Cerebrasqwen-3.8-27bqwen

Quick Info

Powered by
Provider
Cerebras
Model key
qwen-3.8-27b
Release date
Aug 14, 2026
Last updated
Sep 3, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.99
Output token cost
$1.49

Limits

Output tokens
40,960 tokens
Context window
131,072 tokens

Transparent token rates

Compare Qwen3.8 27B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.8 27B

Cerebras

Coverage

Rost Glukhov's Medium article (published August 5/18, 2026) highlights the release of Qwen3.8-27B open weights on Hugging Face (https://huggingface.co/Qwen/Qwen3.8-27B) and ModelScope, alongside Qwen3.8-Max. While Alibaba's Qwen3.8-Max garnered headlines as a 2.4-trillion-parameter flagship with improvements in coding, The article notes that Qwen3.8-Max, despite being open-weight, requires approximately 1.2 terabytes for weights alone at four-bit precision, plus additional space for quantization metadata, runtime buffers, caches, vision components, and distributed-serving overhead, placing it firmly in cluster-mode territory. In cont

Cerebras

CoverageBenchmark

Northflank's August 17, 2026 post provides developer-focused coverage of Qwen3.8-27B, framing it as a 27-billion-parameter open-weight model from Alibaba's Qwen team suited to coding, reasoning, agentic, and multimodal workloads. It states that quantized versions can run on a single GPU including 24GB configurations fo For developers choosing between closed APIs and self-hosted open models, the piece argues Qwen3.8-27B is competitive with much larger models on coding and reasoning benchmarks while remaining practical to run on a single GPU, making it an interesting option for teams that want strong performance without depending entir

Cerebras

CoverageBenchmark

Kingy.ai's launch-day review covers Qwen3.8-27B directly, describing it as a 27.78-billion-parameter dense multimodal model from Alibaba's Qwen team released on August 14, 2026. The article reports the model ships under Apache 2.0, accepts text, images, and video, and has a native 262,144-token context window, and it e The piece concludes Qwen3.8-27B is a leading candidate for the best dense, locally deployable multimodal model around the 30B-parameter mark, combining strong agentic coding, computer-use, and vision-language results with a checkpoint that can realistically be quantized onto a high-end workstation. It cautions that dra

Cerebras

Coverage

An NVIDIA DGX Spark / GB10 forum thread from August 8, 2026, by user "wentbackward" announced that Qwen3.8-27B would arrive the following week alongside full Qwen3.8 open-weights releases, linking to the Qwen blog post at qwen.ai/blog?id=qwen3.8. The thread gained significant traction with over 10.6k views and was link The discussion centers on local deployment scenarios including Qwen3.8-27B-MixedInt4-AutoRound optimized for a single DGX Spark, NVFP4 configurations with vLLM+MTP measurements supporting up to 1M context, and TP=4 serving at 80 tok/s sustained with peaks at 92 tok/s. These community experiments demonstrate the model's

Cerebras

Coverage

Latent Space's AINews weekday roundup from August 3, 2026, covered the Qwen3.8-Max (2.4T) and Qwen3.8-27B open-weights announcements for coding and "cowork" workloads. The newsletter noted that after the "Qwen Exodus" and management changes, there was doubt about whether Qwen would continue releasing relevant open mode Qwen3.8-Max was described as capable of 10+ days of unattended autonomous coding, building a self-evolving coding harness from scratch over a multi-week autonomous run. Additional capabilities highlighted include autonomous AI research (rebuilding a complete paper's pipeline and running an iterative research loop over

Cerebras

CoverageAnalysis

Alibaba's Tongyi Lab released Qwen3.8-27B on August 14, 2026, at 15:00 UTC as a 27.78-billion-parameter dense multimodal language model under Apache 2.0, distributed via Hugging Face. The release is positioned as a significant milestone for locally deployable AI, reportedly achieving competitive performance with models The model architecture features 64 Transformer blocks with a hidden dimension of 5,120, a 24 query head / 4 KV head GQA configuration with head dimension 256, a vocabulary of 248,320 tokens, and a native 262,144-token context window extensible to 1M tokens via YaRN. The hybrid attention mechanism uses a 3:1 ratio: 48 G

Videos about Qwen3.8 27B

More models around Qwen3.8 27B