Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Novita AI logo

Model details

Qwen3.6 35B-A3B

We couldn't load the overview just now. Please try again in a little while.

Novita AIqwen/qwen3.6-35b-a3bqwen

Quick Info

Powered by
Provider
Novita AI
Model key
qwen/qwen3.6-35b-a3b
Release date
Apr 17, 2026
Last updated
Apr 17, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.248
Output token cost
$1.485

Limits

Output tokens
65,536 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen3.6 35B-A3B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.6 35B-A3B

Kilo Gateway

Coverage

Qwen open-sourced Qwen3.6-35B-A3B on April 14, 2026, a sparse Mixture-of-Experts model with 35B total parameters and only 3B active. The release targets agentic coding and supports both multimodal thinking and non-thinking modes. It is available on Qwen Studio, via the Alibaba Cloud Model Studio API as Qwen3.6-Flash, and as open weights on Hugging Face and ModelScope. Benchmarks show the model scoring 73.4 on SWE-bench Verified, 67.2 on SWE-bench Multilingual, 49.5 on SWE-bench Pro, 51.5 on Terminal-Bench 2.0, and 1397 on QwenWebBench, outperforming its Qwen3.5-35B-A3B predecessor and rivaling denser peers like Qwen3.5-27B. The post frames it as one of the most versatile open-source models available, built for repository-level reasoning tasks.

Kilo Gateway

CoverageBenchmark

Millstone AI benchmarked Qwen3.6-35B-A3B in FP8 on a single AMD MI300X with 192GB VRAM using vLLM. Peak throughput reached 255.6 tokens per second at 3 concurrent requests, with 100 percent success rate across 2.4K requests. The May 10, 2026 report measured 3 concurrent users capacity at 32K context. Single-request generation speed ran 113.9 tok/s at 1K context, dropping to 7.3 tok/s at 128K context. TTFT ranged from 116ms at 1K to 27.1s at 128K, with throughput scaling 2.5x from 1 to 3 concurrent requests. Tests covered 1K to 128K token contexts without prompt caching or speculative decoding.

Kilo Gateway

CoverageBenchmark

BenchLM's October 8, 2026 snapshot of Qwen3.6-35B-A3B reports a composite capability score of 43.3 out of 100, ranking 121 of 216 tracked public models. The model shows particularly strong multimodal evidence at rank 37 of 49, with 14 verified benchmarks averaging 51.0 for screenshots, documents, and charts. Category breakdown shows Agentic at 29.2 (rank 86 of 122, 11 benchmarks), Coding at 34.3 (rank 80 of 146, 6 benchmarks), Knowledge at 41.9 (rank 99 of 174, 5 benchmarks), and Math at 70.7 across 5 verified benchmarks. The page lists 41 displayable benchmark rows and 136 tok/s measured speed, with 262K context reported.

Kilo Gateway

Coverage

A third-party guide details Qwen 3.6-35B-A3B as Alibaba's April 15, 2026 open-weight release, a 35B/3B active sparse MoE with 256 experts, Apache 2.0 license, 262K native context extensible to 1M via YaRN, and vision support for image and video. The model gained 540+ Hugging Face likes and 21K+ downloads in its first two days. Reported scores include 73.4% SWE-bench Verified, 49.5% SWE-bench Pro, 67.2% SWE-bench Multilingual, 51.6% Terminal-Bench 2.0, 92.7% AIME 2026, 86.0% GPQA Diamond, and 85.2% MMLU-Pro. The Terminal-Bench score notably beats dense Qwen 3.5-27B by a wide margin, highlighting CLI task handling improvements.

Kilo Gateway

CoverageBenchmark

Millstone AI benchmarked Qwen3.6-35B-A3B in FP8 precision on a single RTX Pro 6000 Blackwell GPU with 96GB VRAM using vLLM. The model hit 449.0 tokens per second peak throughput at 5 concurrent requests, with a 100 percent success rate across tested scenarios. Capacity reached 41 concurrent users at 32K context length. Single-request generation speed measured 196.4 tok/s at 1K context declining to 116.3 tok/s at 256K context, while TTFT ranged from 55ms at 1K to 23.9s at 256K. The May 25, 2026 report tests context lengths from 1K to 256K tokens without prompt caching or speculative decoding, with full-precision KV cache.

Kilo Gateway

Coverage

The Hugging Face model card for Qwen3.6-35B-A3B documents a Causal Language Model with Vision Encoder, 35B total and 3B activated parameters, 40 layers, 256 experts with 8 routed plus 1 shared, and a hidden layout mixing Gated DeltaNet linear attention with Gated Attention. Native context is 262,144 tokens, extensible up to 1,010,000 via YaRN. Key release highlights include improved agentic coding for frontend workflows and repository-level reasoning, plus a new option to retain reasoning context from historical messages for iterative development. Weights are compatible with Hugging Face Transformers, vLLM, SGLang, and KTransformers, with MTP multi-step training support.

OpenRouter

CoverageBenchmark

The OpenRouter listing for qwen/qwen3.6-35b-a3b confirms the model as an open-weight multimodal release from Alibaba Cloud with 35 billion total parameters and 3 billion active per token, using a hybrid sparse MoE that combines Gated DeltaNet linear attention with standard gated attention layers for efficient inference The listing records the model's OpenRouter release date as April 27, 2026, with listed pricing of $0.05 per 1M input tokens and $0.70 per 1M output tokens. OpenRouter also displays a weighted-average effective price based on what customers actually pay across its routed providers, along with a price-history chart spann

OpenRouter

Coverage

On April 2, 2026, Alibaba's Qwen team open-sourced Qwen3.6-35B-A3B as the first open-weight variant of the Qwen3.6 generation, releasing it under Apache 2.0 on Hugging Face alongside the proprietary Qwen3.6-Plus API model. The page frames the release with the tagline "Agentic Coding Power, Now Open to All," noting the Architecturally, Qwen3.6-35B-A3B is a sparse MoE with 256 experts where 8 routed plus 1 shared expert activate per token, yielding 3B active parameters out of 35B total. The 40-layer stack uses a repeating block of three Gated DeltaNet (linear attention) layers followed by one Gated Attention layer, each paired with an

Videos about Qwen3.6 35B-A3B

More models around Qwen3.6 35B-A3B