Currently listed through these providers:
Model details
Qwen3.6 35B-A3B
Qwen 3.6 35B-A3B is a Mixture-of-Experts model from the Qwen family that targets agentic workflows and complex reasoning tasks. According to the Salad AI Gateway overview, it is positioned as the flagship MoE option in the gateway's closed-beta lineup, recommended specifically for agentic scenarios and multi-step reasoning where tool use and structured outputs are central. The naming suggests roughly 35 billion total parameters with about 3 billion active per token, a configuration that typically balances capability against compute cost, and the public FP8 quantization discussion on the NVIDIA developer forums indicates community attention to efficient inference variants suitable for local and edge GPU deployment.
In practical use, the model fits builders who need strong reasoning plus reliable tool calling and structured output, and its open-weights nature means teams can also run or fine-tune it locally. Salad AI Gateway routes requests through an OpenAI-compatible endpoint powered by SaladCloud's distributed GPU network, which keeps setup minimal for developers using popular coding and agent tools, while the gateway still exposes the model's reasoning and tool-calling capabilities at the API level. Community testers on the NVIDIA forums have been actively tuning configurations, signaling an active ecosystem around this release and suggesting the model is being explored for stable agent pipelines where reasoning quality matters more than raw throughput.
Quick Info
Powered by- Provider
- SaladCloud AI Gateway
- Model key
- qwen3.6-35b-a3b
- Release date
- Apr 17, 2026
- Last updated
- Apr 17, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.09
- Output token cost
- $0.60
Limits
- Input tokens
- 262,144 tokens
- Output tokens
- 262,144 tokens
- Context window
- 262,144 tokens