Sulat.com
AI models
InferX logo

Model details

Qwen3.6 35B A3B FP8

Qwen3.6 35B A3B FP8 is the first open-weight variant of the Qwen3.6 series, positioned as a stability-and-utility refresh that learns from community feedback on the earlier Qwen3.5 line. It is a causal language model with an integrated vision encoder, built around a sparse Mixture-of-Experts design totaling roughly 35B parameters while activating only about 3B per token across 256 experts (8 routed plus 1 shared). The underlying architecture weaves together Gated DeltaNet and Gated Attention layers, a hybrid that aims to balance long-range memory with precise local recall, and the model is released under an Apache 2.0 license for broad downstream use.

For deployment, the FP8 checkpoint applies fine-grained quantization with a block size of 128, retaining near-identical performance to the original precision while reducing memory pressure, and the weights ship in a Hugging Face Transformers format compatible with vLLM, SGLang, and KTransformers. The model targets agentic coding and repository-level reasoning with stronger frontend workflow handling, plus a thinking-preservation option that carries reasoning context across messages to keep iterative development sessions coherent. A native 262,144-token context, extendable toward one million tokens via YaRN, paired with multimodal text, image, and video inputs, makes it well suited to long-document analysis, multi-file code agents, and multimodal assistants on a single high-memory accelerator.

InferXQwen3.6-35B-A3B-FP8qwen

Quick Info

Powered by
Provider
InferX
Model key
Qwen3.6-35B-A3B-FP8
Release date
Apr 17, 2026
Last updated
Apr 17, 2026
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
65,536 tokens
Context window
262,000 tokens

Latest news about Qwen3.6 35B A3B FP8

Videos about Qwen3.6 35B A3B FP8

More models around Qwen3.6 35B A3B FP8