Currently listed through these providers:
Model details
Qwen3.6 27B
Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen team at Alibaba and the first open-weight variant of the Qwen3.6 series, arriving after the February release of the Qwen3.5 line. The official model card positions it as a post-trained release in the Hugging Face Transformers format, with artifacts compatible with vLLM, SGLang, and KTransformers, and it is distributed under the Apache 2.0 license. The architecture combines a 5,120-dim hidden layout with 64 layers arranged as a hybrid stack of sixteen Gated DeltaNet blocks (48 V and 16 QK linear-attention heads at head dim 128) followed by one Gated Attention block (24 Q heads and 4 KV heads), giving the model a mix of linear-attention throughput and standard attention precision suited to long context.
The release leans into practical developer experience, emphasizing stability and real-world usefulness over benchmark theater. Two named upgrades stand out: agentic coding, where the model handles frontend workflows and repository-level reasoning with greater fluency, and a new Thinking Preservation option that retains reasoning context across historical messages, reducing overhead during iterative development. Through its hosted endpoints it reaches a 262,144-token context window, supports tool calling, temperature control, and structured output, and serves 201 languages and dialects, making it a strong fit for long-running code agents, multi-step refactors, and mixed-language assistants where preserving the model's internal scratchpad matters as much as raw inference cost.
Quick Info
Powered by- Provider
- SiliconFlow
- Model key
- Qwen/Qwen3.6-27B
- Release date
- Apr 22, 2026
- Last updated
- Apr 22, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.30
- Output token cost
- $3.20
Limits
- Output tokens
- 262,144 tokens
- Context window
- 262,144 tokens