Currently listed through these providers:
Model details
Qwen 3.5 397B
Qwen 3.5 397B is a large-scale foundation model built on an efficient hybrid architecture that combines Gated Delta Networks with a sparse Mixture-of-Experts design. By activating only 17 billion parameters per token out of its 397 billion total, the model achieves high-throughput inference with reduced latency and cost compared to denser, larger-scale alternatives. This design intent focuses on providing enterprise-grade intelligence that remains practical for deployment, allowing developers to leverage advanced reasoning, coding, and complex AI workflows without the need for massive, unmanageable GPU clusters.
The model benefits from a post-training lineage that emphasizes multimodal learning, utilizing early fusion training on multimodal tokens to achieve strong performance across visual understanding and text-based reasoning benchmarks. By integrating breakthroughs in reinforcement learning scale and architectural efficiency, it provides a versatile tool for agents and cross-generational tasks. Its ability to trade blows with trillion-parameter models while maintaining a more accessible footprint makes it a significant development for organizations looking to own and control their AI infrastructure while maintaining competitive performance levels.
Quick Info
Powered by- Provider
- Venice AI
- Model key
- qwen3-5-397b-a17b
- Release date
- Feb 16, 2026
- Last updated
- Jun 11, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.75
- Output token cost
- $4.50
Limits
- Output tokens
- 32,768 tokens
- Context window
- 128,000 tokens
Transparent token rates
Compare Qwen 3.5 397B pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.