Currently listed through these providers:
Model details
Qwen3 235B-A22B
Qwen3-235B-A22B is a flagship open-weight language model from Alibaba Cloud's Qwen team, built around a Mixture-of-Experts architecture that activates roughly 22B of its 235B total parameters per inference pass. This design aims to balance strong reasoning quality with lower-latency generation, and the model is distributed under the Apache 2.0 license with no gating on Hugging Face, making it freely usable for research and production experimentation. It targets text generation as its core modality, with emphasis on instruction-following, agent-based tasks, and multilingual translation across more than 100 languages.
A distinguishing feature of Qwen3-235B-A22B is its switchable operational mode: a "thinking" mode optimized for complex multi-step reasoning and a "non-thinking" mode designed for fast, lightweight dialogue. Users can toggle between these behaviors through the chat template or simple soft-switch prompts such as /think and /no_think, letting applications allocate compute based on task complexity. The model also incorporates extensive pretraining and alignment with human preferences, positioning it as a flexible foundation for assistants, agentic workflows, and multilingual applications where reasoning depth needs to be adjustable on demand.
Quick Info
Powered by- Provider
- Kilo Gateway
- Model key
- qwen/qwen3-235b-a22b
- Release date
- Apr 1, 2025
- Last updated
- Apr 1, 2025
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.455
- Output token cost
- $1.82
Limits
- Output tokens
- 8,192 tokens
- Context window
- 131,072 tokens
Transparent token rates
Compare Qwen3 235B-A22B pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Qwen3 235B-A22B
No articles yet. Fetch the latest news to show it here.