Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Venice AI logo

Model details

Qwen 3.8 27B

Qwen 3.8 27B is a dense, instruction-tuned model with 27 billion parameters, positioned as a versatile system for coding, research, real-world work, vision, and agentic tasks. Its architecture is described as part of the latest Qwen generation, with a broad context suited to work that requires maintaining substantial conversational or project material.

The model is designed for practical local and edge-oriented development as well as general-purpose generation. AMD reports early performance on Ryzen AI Max+ hardware reaching up to 24.5 tokens per second, while its open framework support enables deployment through llama.cpp on systems with more than 24 GB of usable graphics memory. This combination makes it a strong fit for developers experimenting with local AI, multimodal interfaces, and longer-running coding or research workflows.

Venice AIqwen-3-8-27bqwen

Quick Info

Powered by
Provider
Venice AI
Model key
qwen-3-8-27b
Release date
Aug 17, 2026
Last updated
Aug 18, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.45
Output token cost
$3.20

Limits

Output tokens
65,536 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen 3.8 27B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen 3.8 27B

Venice AI

CoverageBenchmark

An NVIDIA Developer Forums post (Aug 22, 2026) by user Ama5u presents a detailed serving-stack benchmark of Qwen3.8-27B on a single DGX Spark (GB10) system. The comparison cross four configurations: vLLM+MTP versus SGLang+DFlash2 speculative decoding, each tested under greedy and Qwen's official "thinking-mode" sampler The reported results show SGLang+DFlash2 with greedy sampling achieving quality 91/100, responsiveness 43, median turn 3.6s, and wall time 929s—substantially faster than vLLM+MTP (quality 90/100, responsiveness 19, wall time 2386s). Decode throughput with DFlash2 was roughly 2.5× faster than vLLM+MTP across code, reaso

Venice AI

CoverageBenchmark

A MindStudio blog post titled 'Qwen 3.8 27B Benchmarked: Agentic Index, Vision, and Reasoning Tests,' dated August 20, 2026, signals an independent benchmark of the Qwen3.8-27B family across agentic, vision, and reasoning evaluations on local inference hardware. The candidate excerpt is dominated by the site's cookie/c Because the substantive content of the article was not captured in the scrape, this entry is accepted as a pointer to on-topic Qwen 3.8 27B benchmark coverage rather than as a source of verifiable numbers. MindStudio is a third-party local-AI tooling vendor, so its results should be read with awareness of commercial in

Venice AI

Coverage

A third-party Medium article by Rost Glukhov reports on the open-weight release of Alibaba's Qwen3.8-27B, citing an official Qwen account announcement that the 27B variant would be made open-weight alongside Qwen3.8-Max. The page excerpt confirms the weights are hosted on Hugging Face at huggingface.co/Qwen/Qwen3.8-27B The same excerpt describes Qwen3.8-27B as benefiting from improvements in coding, professional "cowork" tasks, full-stack development, data analysis, and office workflows relative to prior Qwen generations. The author argues that while Qwen3.8-Max draws headline attention due to its scale, the 27B open-weight variant i

Venice AI

CoverageAnalysis

A Hacker News thread (item 49334544, dated Aug 18, 2026, 381 points) reports that Qwen3.8 27B scored 52 on the Artificial Analysis benchmark, compared to 38 for Qwen3.6 27B. According to commenters, the new score beats all medium-tier models (40B–150B) and ties DeepSeek V4 Flash 0731, which ranks 5th in the large-model Thread participants note that the high benchmark score likely comes with substantially increased token consumption—approximately 2.3× GPT Luna Max and nearly 2× Kimi K3—attributed to Max reasoning mode producing very long reasoning traces. Discussion also covers tokens-per-second comparisons between Qwen and Gemma 4 mo

Venice AI

CoverageBenchmark

A HackerNoon post describes Qwen3.8-27B-Cold-Fusion-GAIN-V1.1, a 27-billion-parameter instruction-tuned community build by "DavidAU" that applies the COLD FUSION training methodology, combining an internally developed GAIN technique with Unsloth's training infrastructure, to compress thinking tokens to roughly one-tent Built on Qwen's 3.8 architecture, the model runs on the transformers library and targets customer support automation, code review systems, and real-time data analysis where complex problem-solving output is needed but long internal reasoning is not affordable. The article itself is flagged as AI-assisted and includes p

Venice AI

Coverage

A Hacker News discussion thread (item 49150809) references the Qwen3.8-27B open-weight release scheduled for the week following the Qwen3.8-Max announcement. Commenters note that the prior Qwen3.6-27B was widely regarded as one of the best local models in its size class, and express optimism that Qwen3.8-27B will impro Community comments in the same thread share practical local-deployment experiences with earlier Qwen3.6 variants (27B and 35B-A3B) running on consumer hardware including RTX 5090s, AMD R9700s, and Apple Silicon Macs, using harnesses such as OpenCode for agentic coding workflows. Users describe the Qwen3.6 line as a "wo

Venice AI

CoverageBenchmark

A third-party guide from Atomic Chat details how to run Alibaba's Qwen 3.8 27B locally, describing it as an open-weight dense 27B multimodal language model from the Qwen 3.8 family that fits on a 24 GB GPU or a 32 GB Mac at 4-bit quantization. The article cites an August 14, 2026 release of the 27B weights alongside th The guide also walks through Atomic Dynamic GGUF builds, the hardware needed to run the model, and concrete setup steps for both Atomic Chat and llama.cpp. It references sibling resources such as a Qwen 3.8 27B uncensored variant and a broader Qwen-family local-inference guide, framing the model as a locally deployable

Videos about Qwen 3.8 27B

More models around Qwen 3.8 27B