Currently listed through these providers:
Model details
Qwen3 32B
Qwen3-32B is part of the latest generation of Qwen series models, a family that pairs dense architectures with mixture-of-experts variants. The model card positions it as a causal language model trained through both pretraining and post-training stages, with 32.8B total parameters of which 31.2B are non-embedding, organized across 64 layers. Attention uses grouped-query design with 64 query heads and 8 key-value heads, and the native context length is 32,768 tokens, providing a workable window for extended reasoning chains, multi-document prompts, and longer agent traces without immediately invoking extended-context techniques.
A defining practical feature is the ability to switch seamlessly between a thinking mode, aimed at complex logical reasoning, mathematics, and code generation, and a non-thinking mode optimized for efficient general-purpose dialogue, all within a single model. The release emphasizes that reasoning gains surpass the prior QwQ line in thinking mode and Qwen2.5 instruct in non-thinking mode across math, code, and commonsense reasoning, alongside stronger human preference alignment for creative writing, role-play, and instruction following. Agent capabilities are highlighted as a specialty, with precise tool integration available in both modes, making the model a reasonable fit for developers who want one open-weight checkpoint that can toggle between deep step-by-step problem solving and lighter conversational use, and that supports multilingual workloads across 100+ languages and dialects.
Quick Info
Powered by- Provider
- Hugging Face
- Model key
- Qwen/Qwen3-32B
- Release date
- Apr 1, 2025
- Last updated
- Apr 1, 2025
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.29
- Output token cost
- $0.59
Limits
- Output tokens
- 16,384 tokens
- Context window
- 131,072 tokens