Sulat.com
AI models
ai& logo

Model details

Qwen3.6 27B

Qwen3.6-27B is the first open-weight variant released in the Qwen3.6 series, positioned by its creators as a stability-focused post-trained model for real-world developer work. It ships as a Causal Language Model paired with a vision encoder, distributed in Hugging Face Transformers format with weights and configuration files published openly. The release artifacts are stated to be compatible with Hugging Face Transformers, vLLM, SGLang, and KTransformers, giving developers flexibility in choosing their inference stack.

The model targets agentic coding and iterative development, with creators highlighting stronger handling of frontend workflows and repository-level reasoning alongside a new Thinking Preservation option that retains reasoning context across historical messages. Its hybrid hidden layout interleaves Gated DeltaNet linear attention blocks with standard Gated Attention, spanning 64 layers around a 27B-parameter backbone, a design aimed at balancing long-context throughput with precise local reasoning. Practically, it suits developers who want a self-hostable coding assistant with open weights, structured reasoning support, and a large context window for working across substantial codebases.

ai&qwen/qwen3.6-27bqwen

Quick Info

Powered by
Provider
ai&
Model key
qwen/qwen3.6-27b
Release date
Apr 22, 2026
Last updated
Apr 22, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.32
Output token cost
$3.20

Limits

Output tokens
65,536 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen3.6 27B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.6 27B

Vultr

CoverageBenchmark

A third-party technical deep dive published on kie.ai on July 14, 2026 documents what independent testers have measured on Qwen 3.6 27B roughly three weeks after the weights landed on Hugging Face. The model is described in the supplied excerpt as a 27-billion-parameter dense transformer with a hybrid linear-attention The deep dive reports concrete developer-relevant numbers: NVFP4 quantization hits MMLU accuracy of 0.8446 and delivers roughly 2.6–2.86× decode speedup over BF16 in a vLLM benchmark. On DGX Spark hardware, community testers measured 28–33 tokens/second single-session throughput on the NVFP4 build via vLLM 0.24.0, whil

Videos about Qwen3.6 27B

More models around Qwen3.6 27B