Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Deep Infra logo

Model details

Qwen3.6 35B A3B

As a member of the Qwen family from Alibaba, this 35B-parameter Mixture-of-Experts model (designated A3B) continues the lineage of open-weight releases that have made Qwen a practical choice for developers building agentic and reasoning-driven applications. The Qwen3.6 iteration gained rapid community attention shortly after its April 2026 release, with discussions surfacing on the NVIDIA DGX Spark / GB10 forums about both the base model and an FP8 quantized variant optimized for efficient deployment on consumer and workstation GPUs. Third-party commentary published in Towards AI framed the release as a step toward addressing persistent context-retention challenges in long-running AI agent workflows.

On Deep Infra, the model is served via an OpenAI-compatible endpoint, which simplifies integration into existing toolchains and agent frameworks. Independent pricing trackers confirm that Deep Infra currently offers the most competitive input-token rate among major inference providers for this model, undercutting alternatives such as OpenRouter, IO.NET, and Scaleway. This combination of a MoE architecture, quantization-friendly variants, and cost-effective hosted serving makes the model a reasonable fit for teams building production agent systems, multi-step reasoning pipelines, and structured-output workflows where both throughput and per-token economics matter.

Deep InfraQwen/Qwen3.6-35B-A3Bqwen

Quick Info

Powered by
Provider
Deep Infra
Model key
Qwen/Qwen3.6-35B-A3B
Release date
Apr 1, 2026
Last updated
Apr 1, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.10
Output token cost
$0.95

Limits

Output tokens
81,920 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen3.6 35B A3B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.6 35B A3B

Deep Infra

Coverage

A third-party pricing tracker confirms DeepInfra's standard rate for Qwen3.6 35B A3B at $0.100 per 1M input tokens and $0.950 per 1M output tokens, last updated August 23, 2026, and published August 11, 2026. The model is listed as a text-only model created by Alibaba, deployed alongside an OpenAI-compatible endpoint o Across the four providers tracked, DeepInfra currently offers the lowest input price for Qwen3.6 35B A3B, compared with OpenRouter at $0.140, IO.NET at $0.187, and Scaleway at $0.292 per 1M input tokens. The tracker notes one standard pricing mode without batch or cached-input rates, helping developers benchmark cross-

Merge Gateway

Coverage

On April 2, 2026, Alibaba's Qwen team open-sourced Qwen3.6-35B-A3B under Apache 2.0 on Hugging Face as the first open-weight variant of the Qwen3.6 generation, shipped alongside the proprietary Qwen3.6-Plus API model. The 35-billion-parameter Mixture-of-Experts model activates just 3B parameters per token and was frame Qwen3.6-35B-A3B is a sparse MoE with 256 experts, of which 8 routed plus 1 shared activate per token, using a 40-layer stack of three Gated DeltaNet (linear attention) layers followed by one Gated Attention layer, each paired with an MoE feed-forward block; the hidden dimension is 2048 and expert intermediate dimension

Videos about Qwen3.6 35B A3B

More models around Qwen3.6 35B A3B