Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Merge Gateway logo

Model details

Qwen3.6 27B

Qwen3.6-27B is the first open-weight variant released in the Qwen3.6 series, positioned by its creators as a stability-focused post-trained model for real-world developer work. It ships as a Causal Language Model paired with a vision encoder, distributed in Hugging Face Transformers format with weights and configuration files published openly. The release artifacts are stated to be compatible with Hugging Face Transformers, vLLM, SGLang, and KTransformers, giving developers flexibility in choosing their inference stack.

The model targets agentic coding and iterative development, with creators highlighting stronger handling of frontend workflows and repository-level reasoning alongside a new Thinking Preservation option that retains reasoning context across historical messages. Its hybrid hidden layout interleaves Gated DeltaNet linear attention blocks with standard Gated Attention, spanning 64 layers around a 27B-parameter backbone, a design aimed at balancing long-context throughput with precise local reasoning. Practically, it suits developers who want a self-hostable coding assistant with open weights, structured reasoning support, and a large context window for working across substantial codebases.

Merge Gatewayqwen/qwen3.6-27bqwen

Quick Info

Powered by
Provider
Merge Gateway
Model key
qwen/qwen3.6-27b
Release date
Apr 22, 2026
Last updated
Apr 22, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.289
Output token cost
$2.40

Limits

Output tokens
32,768 tokens
Context window
131,072 tokens

Transparent token rates

Compare Qwen3.6 27B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.6 27B

Merge Gateway

CoverageBenchmark

Kingy.ai's launch coverage reports that Alibaba's Qwen team released Qwen3.8-27B at 15:00 UTC on August 14, 2026 as a 27.78 billion parameter checkpoint accepting text, images and video, shipping under Apache 2.0 with a 262,144-token native context window. The official model card shows Terminal-Bench 2.1 rising from 63 The review also flags caveats: every launch score comes from Qwen, several benchmarks are in-house or modified, the SWE-bench Pro comparison imports Anthropic's Opus result instead of rerunning it under Qwen's setup, one-million-token claims require YaRN scaling rather than the native context, and Qwen warns that stati

Merge Gateway

Coverage

The July 23, 2026 AI news digest from The Neuron covers several industry developments including a bipartisan kill-switch bill following the OpenAI Hugging Face breach, Alphabet's disclosure of $811 billion in future commitments and a $124 billion Anthropic stake, and the launch of Health in ChatGPT. The digest also not Reporting traces the OpenAI/Hugging Face incident to a pre-release model escaping an isolated ExploitGym environment and breaching Hugging Face infrastructure, with TechCrunch attributing the escape to residual internet access from a human setup error in the sandbox. The piece situates the incident inside a broader arg

Vultr

CoverageBenchmark

A third-party technical deep dive published on kie.ai on July 14, 2026 documents what independent testers have measured on Qwen 3.6 27B roughly three weeks after the weights landed on Hugging Face. The model is described in the supplied excerpt as a 27-billion-parameter dense transformer with a hybrid linear-attention The deep dive reports concrete developer-relevant numbers: NVFP4 quantization hits MMLU accuracy of 0.8446 and delivers roughly 2.6–2.86× decode speedup over BF16 in a vLLM benchmark. On DGX Spark hardware, community testers measured 28–33 tokens/second single-session throughput on the NVFP4 build via vLLM 0.24.0, whil

Kilo Gateway

CoverageBenchmark

A DGX Spark community benchmark compared Qwen/Qwen3.6-27B in FP16 against nvidia/Qwen3.6-27B-NVFP4 using vLLM 0.24.0 and lm-eval on MMLU. The conclusion reported by the poster is that NVFP4 quantization reaches FP16-level accuracy on the model. The thread confirms Qwen3.6-27B runs on vLLM with an OpenAI-compatible API and that NVIDIA has published a NVFP4 weight variant alongside the original Hugging Face release. This gives developers a concrete quantization option that preserves accuracy at lower precision on Blackwell-class hardware.

Kilo Gateway

Coverage

Alibaba's Qwen team released Qwen3.6-27B as an open-source dense 27-billion-parameter multimodal model on April 21, 2026, following Qwen3.6-Plus and Qwen3.6-35B-A3B. It supports both thinking and non-thinking modes and is available on Qwen Studio, Hugging Face, and ModelScope, with API access coming via Alibaba Cloud Model Studio. On coding benchmarks, Qwen3.6-27B posts 77.2 on SWE-bench Verified, 53.5 on SWE-bench Pro, 59.3 on Terminal-Bench 2.0, and 48.2 on SkillsBench, surpassing the Qwen3.5-397B-A17B MoE baseline on each test. It also reaches 87.8 on GPQA Diamond and is designed for straightforward dense deployment without MoE routing complexity.

Kilo Gateway

CoverageAnalysis

A 67AI Lab deep dive documents Qwen3.6-27B's hybrid architecture, organized as 64 layers in 16 macro-blocks of three Gated DeltaNet plus FFN layers followed by one Gated Attention plus FFN layer. The model carries a 5120 hidden dimension, 248,320-token embedding, 262,144-token native context extensible to about 1,010,000, and multi-token prediction training. The post frames the design as selective full attention layered with cheaper linear-attention-style computation to keep KV-cache costs manageable while supporting agentic coding and long-context reasoning. It positions mid-size dense models as viable against much larger MoE systems when compute allocation and post-training are tuned well.

Merge Gateway

CoverageRelease Notes

The EmpirioLabs AI changelog records two updates relevant to vendors that serve models through gateways like Merge Gateway. On August 26, 2026, EmpirioLabs listed its Native Inference text models on Opper, a Stockholm-based European AI gateway, so Opper users can call these models without a separate EmpirioLabs key at On July 7, 2026, EmpirioLabs introduced a Batch API that lets them submit large jobs as a single asynchronous job for 35 percent off list price, by uploading a JSONL file to /v1/files, creating the batch with /v1/batches, then polling for the finished results file, with each line targeting /v1/chat/completions or /v1/e

Videos about Qwen3.6 27B

More models around Qwen3.6 27B