Currently listed through these providers:
Model details
Qwen3.5 397B-A17B FP8
Qwen3.5-397B-A17B is a large sparse Mixture-of-Experts vision-language model released by the QwenLM team as part of the Qwen 3.5 family. It is positioned as a unified multimodal foundation that performs early fusion training on multimodal tokens, aiming to deliver cross-generational reasoning, coding, agent, and visual understanding capabilities that the Qwen team describes as outperforming their earlier Qwen3-VL line on key benchmarks. The base architecture combines expert routing for high-throughput inference with reinforcement learning scaled across million-agent environments, and the family documentation notes support across 201 languages for global developer reach.
This FP8 artifact is the post-trained checkpoint quantized with fine-grained FP8 at a block size of 128, a scheme the model card states preserves performance metrics nearly identical to the original full-precision model. The open weights ship in the Hugging Face Transformers format and are explicitly compatible with serving stacks such as vLLM, SGLang, and KTransformers, making the model well suited to self-hosted deployments that need multimodal inputs and text generation with tool use and structured outputs. The Qwen team clearly separates these open weights from the hosted Qwen3.5-Plus service offered through Alibaba Cloud Model Studio, which adds production extras like a 1M context window and adaptive built-in tools, so practitioners choosing the FP8 release should expect the base feature set rather than the managed platform's enhancements.
Quick Info
Powered by- Provider
- Infomaniak
- Model key
- Qwen/Qwen3.5-397B-A17B-FP8
- Release date
- Feb 15, 2026
- Last updated
- Aug 1, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.99
- Output token cost
- $4.46
Limits
- Input tokens
- 200,000 tokens
- Output tokens
- 65,536 tokens
- Context window
- 200,000 tokens
Latest news about Qwen3.5 397B-A17B FP8
No articles yet. Fetch the latest news to show it here.