Currently listed through these providers:
Model details
Qwen3 30B A3b fp8
Qwen3 30B A3B fp8 is a Mixture-of-Experts text generation model from the Qwen3 family, with approximately 30.53B total parameters distributed in Transformers/Safetensors/PyTorch form and released under the Apache-2.0 license. It is positioned as part of Qwen3's dual-mode design, allowing seamless switching between a thinking mode aimed at complex reasoning, mathematics, and code, and a non-thinking mode optimized for efficient general-purpose dialogue. The FP8 quantization reduces memory footprint relative to the dense Qwen3 variants while keeping weights openly downloadable for self-hosting and experimentation.
The model's open-weights distribution and MoE architecture make it well suited for teams that want to run a mid-sized reasoning-capable language model on their own infrastructure, and independent deployment guides already document serving it with stacks such as SGLang on Kubernetes-based platforms. Because it inherits the broader Qwen3 advances in instruction following, agent-style tool integration, and multilingual support, it fits practical use cases ranging from assistant-style chat and creative writing to structured reasoning tasks, where developers can pick the thinking mode for harder prompts and the lighter non-thinking mode for routine interactions.
Quick Info
Powered by- Provider
- Cloudflare Workers AI
- Model key
- @cf/qwen/qwen3-30b-a3b-fp8
- Release date
- Apr 28, 2025
- Last updated
- Apr 28, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.0509
- Output token cost
- $0.335
Limits
- Output tokens
- 32,768 tokens
- Context window
- 32,768 tokens