Currently listed through these providers:
Model details
Qwen3 32B
Qwen3-32B is a dense 32B-parameter large language model in the Qwen3 family developed by Alibaba, released as an open-weights model for both research and production use. It uses Grouped-Query Attention and is designed to handle extended inputs, with deployment documentation describing a context length of up to 128k tokens that can be stretched to roughly 131k through YaRN extension. The model goes through both pretraining and post-training stages and is intended as a general-purpose foundation model that can serve as a single base for chat assistants, code helpers, and tool-using agents.
A defining feature of Qwen3-32B is its hybrid operation, allowing seamless switching between a thinking mode aimed at complex logical reasoning, mathematics, and coding, and a non-thinking mode optimized for efficient, general-purpose dialogue. The model card describes notable gains over earlier Qwen generations in instruction following, reasoning, text comprehension, mathematics, science, coding, and tool usage, with stronger human-preference alignment for creative writing, role-playing, and multi-turn conversation. It also targets strong agent capabilities and broad multilingual coverage, making it a flexible fit for builders who want one open-weights checkpoint that can alternate between fast dialogue and deeper step-by-step problem solving.
Quick Info
Powered by- Provider
- Kilo Gateway
- Model key
- qwen/qwen3-32b
- Release date
- Apr 1, 2025
- Last updated
- Apr 1, 2025
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.08
- Output token cost
- $0.28
Limits
- Output tokens
- 16,384 tokens
- Context window
- 40,960 tokens
Transparent token rates
Compare Qwen3 32B pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Qwen3 32B
No articles yet. Fetch the latest news to show it here.