Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Abacus logo

Model details

Qwen3.6 27B

Qwen3.6-27B is the first open-weight variant released in the Qwen3.6 series, positioned by its creators as a stability-focused post-trained model for real-world developer work. It ships as a Causal Language Model paired with a vision encoder, distributed in Hugging Face Transformers format with weights and configuration files published openly. The release artifacts are stated to be compatible with Hugging Face Transformers, vLLM, SGLang, and KTransformers, giving developers flexibility in choosing their inference stack.

The model targets agentic coding and iterative development, with creators highlighting stronger handling of frontend workflows and repository-level reasoning alongside a new Thinking Preservation option that retains reasoning context across historical messages. Its hybrid hidden layout interleaves Gated DeltaNet linear attention blocks with standard Gated Attention, spanning 64 layers around a 27B-parameter backbone, a design aimed at balancing long-context throughput with precise local reasoning. Practically, it suits developers who want a self-hostable coding assistant with open weights, structured reasoning support, and a large context window for working across substantial codebases.

AbacusQwen/Qwen3.6-27Bqwen

Quick Info

Powered by
Provider
Abacus
Model key
Qwen/Qwen3.6-27B
Release date
Apr 22, 2026
Last updated
Apr 22, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.32
Output token cost
$3.20

Limits

Output tokens
8,192 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen3.6 27B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.6 27B

Vultr

CoverageBenchmark

A third-party technical deep dive published on kie.ai on July 14, 2026 documents what independent testers have measured on Qwen 3.6 27B roughly three weeks after the weights landed on Hugging Face. The model is described in the supplied excerpt as a 27-billion-parameter dense transformer with a hybrid linear-attention The deep dive reports concrete developer-relevant numbers: NVFP4 quantization hits MMLU accuracy of 0.8446 and delivers roughly 2.6–2.86× decode speedup over BF16 in a vLLM benchmark. On DGX Spark hardware, community testers measured 28–33 tokens/second single-session throughput on the NVFP4 build via vLLM 0.24.0, whil

Kilo Gateway

CoverageBenchmark

A DGX Spark community benchmark compared Qwen/Qwen3.6-27B in FP16 against nvidia/Qwen3.6-27B-NVFP4 using vLLM 0.24.0 and lm-eval on MMLU. The conclusion reported by the poster is that NVFP4 quantization reaches FP16-level accuracy on the model. The thread confirms Qwen3.6-27B runs on vLLM with an OpenAI-compatible API and that NVIDIA has published a NVFP4 weight variant alongside the original Hugging Face release. This gives developers a concrete quantization option that preserves accuracy at lower precision on Blackwell-class hardware.

Kilo Gateway

Coverage

Alibaba's Qwen team released Qwen3.6-27B as an open-source dense 27-billion-parameter multimodal model on April 21, 2026, following Qwen3.6-Plus and Qwen3.6-35B-A3B. It supports both thinking and non-thinking modes and is available on Qwen Studio, Hugging Face, and ModelScope, with API access coming via Alibaba Cloud Model Studio. On coding benchmarks, Qwen3.6-27B posts 77.2 on SWE-bench Verified, 53.5 on SWE-bench Pro, 59.3 on Terminal-Bench 2.0, and 48.2 on SkillsBench, surpassing the Qwen3.5-397B-A17B MoE baseline on each test. It also reaches 87.8 on GPQA Diamond and is designed for straightforward dense deployment without MoE routing complexity.

Kilo Gateway

CoverageAnalysis

A 67AI Lab deep dive documents Qwen3.6-27B's hybrid architecture, organized as 64 layers in 16 macro-blocks of three Gated DeltaNet plus FFN layers followed by one Gated Attention plus FFN layer. The model carries a 5120 hidden dimension, 248,320-token embedding, 262,144-token native context extensible to about 1,010,000, and multi-token prediction training. The post frames the design as selective full attention layered with cheaper linear-attention-style computation to keep KV-cache costs manageable while supporting agentic coding and long-context reasoning. It positions mid-size dense models as viable against much larger MoE systems when compute allocation and post-training are tuned well.

Videos about Qwen3.6 27B

More models around Qwen3.6 27B