Currently listed through these providers:
Model details
Qwen3.6 35B A3B FP8
Qwen3.6 35B A3B FP8 is the first open-weight variant of the Qwen3.6 series, positioned as a stability-and-utility refresh that learns from community feedback on the earlier Qwen3.5 line. It is a causal language model with an integrated vision encoder, built around a sparse Mixture-of-Experts design totaling roughly 35B parameters while activating only about 3B per token across 256 experts (8 routed plus 1 shared). The underlying architecture weaves together Gated DeltaNet and Gated Attention layers, a hybrid that aims to balance long-range memory with precise local recall, and the model is released under an Apache 2.0 license for broad downstream use.
For deployment, the FP8 checkpoint applies fine-grained quantization with a block size of 128, retaining near-identical performance to the original precision while reducing memory pressure, and the weights ship in a Hugging Face Transformers format compatible with vLLM, SGLang, and KTransformers. The model targets agentic coding and repository-level reasoning with stronger frontend workflow handling, plus a thinking-preservation option that carries reasoning context across messages to keep iterative development sessions coherent. A native 262,144-token context, extendable toward one million tokens via YaRN, paired with multimodal text, image, and video inputs, makes it well suited to long-document analysis, multi-file code agents, and multimodal assistants on a single high-memory accelerator.
Quick Info
Powered by- Provider
- InferX
- Model key
- Qwen3.6-35B-A3B-FP8
- Release date
- Apr 17, 2026
- Last updated
- Apr 17, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 65,536 tokens
- Context window
- 262,000 tokens