Currently listed through these providers:
Model details
Qwen3 32B
Qwen3-32B is a dense causal language model built around a 64-layer transformer backbone with grouped query attention—64 heads for queries and 8 for key-value projections. Designed to handle both deep reasoning and responsive dialogue within a single architecture, the model uniquely lets users toggle between thinking mode, which engages step-by-step logical chains suited to mathematics and coding, and non-thinking mode for faster general-purpose interactions. It natively supports a 32K-token context window, enabling longer document reasoning and multi-turn conversations without aggressive compression. The 32.8 billion parameter scale positions it as the flagship dense checkpoint within the broader Qwen3 family, which also includes Mixture-of-Experts variants, reflecting a deliberate design choice to offer a powerful dense option alongside more parameter-heavy architectures.
The Qwen3-32B checkpoint available through Cortecs reflects the post-trained alignment variant of the base pretrained model, refined through standard fine-tuning and preference alignment workflows to strengthen instruction-following, creative writing, and role-playing interactions. Its agent capabilities are particularly noteworthy—the model can integrate external tools in both thinking and non-thinking modes and achieves competitive standing among open-source releases on complex agent-based evaluations. Support for over 100 languages and dialects broadens its applicability across multilingual use cases. The model ships as open weights under the Apache 2.0 license, giving developers and researchers full access to deploy it across diverse infrastructure. Qwen3-32B strikes a practical balance: it is compact enough for single-GPU deployments while delivering reasoning and alignment performance that makes it suitable as a general-purpose foundation for chatbots, coding assistants, and production agents alike.
Quick Info
Powered by- Provider
- Cortecs
- Model key
- qwen3-32b
- Release date
- Apr 1, 2025
- Last updated
- Apr 1, 2025
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.179
- Output token cost
- $0.697
Limits
- Output tokens
- 16,384 tokens
- Context window
- 16,384 tokens
Transparent token rates
Compare Qwen3 32B pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.