Currently listed through these providers:
Model details
Qwq 32B
QwQ 32B belongs to the Qwen series as its dedicated reasoning model, designed to move beyond conventional instruction-following by generating explicit thought processes before delivering answers. This architectural intent shows up in its performance on hard problems, where the model can break down multi-step logic rather than jumping to conclusions. The model leverages a transformer backbone built with RoPE positional encoding, SwiGLU activations, RMSNorm, and grouped query attention spanning 40 query heads and 8 key-value heads across 64 layers. With roughly 32.5 billion parameters—about 31 billion excluding embedding overhead—it sits in a compact middle ground that Alibaba Cloud positions as capable of matching the performance of much larger cutting-edge reasoning systems.
The training pipeline combines pretraining with post-training supervised finetuning and reinforcement learning, a combination that shapes the model's ability to articulate reasoning steps. Its native context window extends to 131,072 tokens, with YaRN fine-tuning recommended when prompts exceed 8,192 tokens to maintain coherence over longer spans. As an open-weights release, QwQ 32B reflects a broader movement toward democratizing advanced reasoning rather than keeping it locked behind proprietary APIs. Its strengths in mathematics and coding make it particularly suited for developers and researchers tackling structured problem domains, while its open availability encourages community experimentation and fine-tuning for specialized reasoning tasks.
Quick Info
Powered by- Provider
- Cloudflare Workers AI
- Model key
- @cf/qwen/qwq-32b
- Release date
- Mar 5, 2025
- Last updated
- Mar 5, 2025
- Knowledge cutoff
- 2024-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.66
- Output token cost
- $1.00
Limits
- Output tokens
- 24,000 tokens
- Context window
- 24,000 tokens