Currently listed through these providers:
Model details
Qwen3 235B A22B FP8
Qwen3 235B A22B FP8 is a mixture-of-experts model with 235 billion total parameters and 22 billion activated per token, built as part of the Qwen series' latest generation to offer both dense and MoE architectures. The model uniquely supports seamless switching between thinking mode for complex logical reasoning, mathematics, and coding, and non-thinking mode for efficient general-purpose dialogue, all within a single model. Its FP8 quantization brings the memory footprint down while preserving nearly identical quality, making it practical for deployment at scale. The architecture is designed for both inference and non-inference modes, enabling the kind of flexible deployment that modern AI applications demand.
The model represents a major leap over its predecessors, with reasoning capabilities that surpass the earlier QwQ series in thinking mode and general capabilities that exceed Qwen2.5-72B-Instruct in non-thinking mode. Training involved extensive pretraining followed by post-training to build strong instruction-following and agent capabilities, supporting precise integration with external tools across both modes. Qwen3 delivers state-of-the-art performance among models of comparable scale, with strong multilingual proficiency spanning over 100 languages and dialects. Optimized for high-throughput, cost-efficient inference and distillation workflows, this model serves teams building reasoning-heavy applications, multilingual products, and agent pipelines that require reliable tool-calling and structured outputs.
Quick Info
Powered by- Provider
- DevPass (LLM Gateway)
- Model key
- qwen3-235b-a22b-fp8
- Release date
- Apr 28, 2025
- Last updated
- Apr 28, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.20
- Output token cost
- $0.80
Limits
- Output tokens
- 8,192 tokens
- Context window
- 40,960 tokens
Latest news about Qwen3 235B A22B FP8
No articles yet. Fetch the latest news to show it here.