Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Blue Claw logo

Model details

Qwen3.6 35B A3B FP8

Qwen3.6 35B A3B FP8 is a fine-grained quantized causal language model built around a Mixture-of-Experts design: roughly 35 billion parameters are present, while about 3 billion are activated per forward pass. Its architecture combines Gated DeltaNet and Gated Attention layers, and the FP8 checkpoint uses a block size of 128 while its reported performance is nearly identical to the original model. The repository provides weights and configuration files for both pre-training and post-training stages, making this variant practical for experimentation in Hugging Face Transformers and compatible serving environments.

The model is positioned for agentic coding, including frontend workflows and repository-level reasoning, with toggleable thinking and an option to preserve reasoning context across turns. Its native context is reported as 262,144 tokens, with YaRN-based extension up to 1,010,000 tokens, so it is especially suited to iterative development, large codebases, and other tasks that benefit from sustained context. The Apache 2.0 licensing and broad framework compatibility also support local deployment, evaluation, and integration into developer tooling.

Blue ClawQwen/Qwen3.6-35B-A3B-FP8qwenbeta

Quick Info

Powered by
Provider
Blue Claw
Model key
Qwen/Qwen3.6-35B-A3B-FP8
Release date
Apr 17, 2026
Last updated
Apr 17, 2026
Input modalities
Output modalities
Capabilities

Limits

Output tokens
65,536 tokens
Context window
131,072 tokens

Latest news about Qwen3.6 35B A3B FP8

Hetzner

CoverageBenchmark

Dre Dyson's hands-on deployment guide focuses on serving Qwen3.6-35B-A3B-FP8 via vLLM, comparing configurations against a Qwen3.5 baseline. The author rebuilt vLLM from dev wheels (0.19.1rc1) using FlashInfer as the attention backend and the Qwen/Qwen3.6-35B-A3B-FP8 checkpoint on a dual-NVIDIA-GPU rig. Peak throughput The article highlights two model-level improvements called out in the Qwen3.6 release notes: agentic coding enhancements for frontend workflows and repository-level reasoning, and a thinking-preservation feature that retains reasoning context from prior conversation turns. It emphasizes the importance of running an est

Blue Claw

CoverageBenchmark

This third-party blog from dredyson.com is a practitioner's hands-on retrospective on running Qwen3.6-35B-A3B-FP8, the same model served via Blue Claw's API. The author describes testing on NVIDIA DGX Spark hardware and outlines the three release features that matter to developers: improved agentic coding for frontend While anecdotal rather than authoritative, the article adds useful real-world context for engineers evaluating the Qwen3.6-35B-A3B-FP8 deployment on Blue Claw, especially for agentic coding pipelines where iterative context retention matters. It is not an official Qwen release note or Blue Claw changelog, so feature cl

Hetzner

CoverageBenchmark

Millstone AI's inference benchmark page documents Qwen3.6-35B-A3B-FP8 as a 35B-parameter Mixture-of-Experts causal language model with 3B activated parameters per token, organized as 256 experts (8 routed plus 1 shared active per forward pass). The page confirms the FP8 quantization uses fine-grained block-128 quantiza The page highlights agentic coding, repository-level reasoning, and frontend workflow improvements as core focus areas, alongside a toggleable thinking mode and a thinking-preservation feature that retains reasoning traces across conversation turns. Reported hardware benchmarks include 256 tok/s peak throughput on a si

Videos about Qwen3.6 35B A3B FP8

More models around Qwen3.6 35B A3B FP8