Currently listed through these providers:
Model details
Qwen3.6 35B A3B FP8
Qwen3.6-35B-A3B-FP8 is the first open-weight variant in the Qwen3.6 family, designed as a multimodal agentic coding model that prioritizes stability and real-world developer utility following community feedback. It is a Mixture-of-Experts architecture with 35 billion total parameters but only 3 billion activated per token, paired with a hybrid attention design that blends Gated DeltaNet linear attention with gated attention layers for efficient long-context inference. Vision input is supported alongside text, and the model ships with native 262K context that can be extended up to roughly 1M tokens for repository-scale reasoning. The FP8 quantization, fine-grained with a block size of 128, keeps runtime behavior nearly indistinguishable from the original while enabling deployment through Hugging Face Transformers, vLLM, SGLang, and KTransformers.
The release emphasizes two practical advances for iterative development: stronger agentic coding, with smoother handling of frontend workflows and repository-level reasoning, and a new thinking-preservation option that retains reasoning context from prior messages to reduce overhead across multi-turn sessions. Together AI's published numbers place it at 73.4 percent on SWE-bench Verified, 80.4 percent on LiveCodeBench v6, and 51.5 percent on Terminal-Bench 2.0, while multimodal evaluations reach 81.7 percent on MMMU and 86.4 percent on MathVista mini for chart, document, and real-world image understanding. These traits make the model a good fit for teams building coding assistants and tool-using agents that need long-context memory, vision-aware reasoning, and the cost efficiency of sparse activation served in FP8.
Quick Info
Powered by- Provider
- Hetzner
- Model key
- Qwen/Qwen3.6-35B-A3B-FP8
- Release date
- Apr 17, 2026
- Last updated
- Apr 17, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 262,144 tokens
- Context window
- 262,144 tokens
Latest news about Qwen3.6 35B A3B FP8
No articles yet. Fetch the latest news to show it here.