Currently listed through these providers:
Model details
Qwen: Qwen3 235B A22B Thinking 2507 (retires Oct 9)
Qwen3-235B-A22B-Thinking-2507 is a causal language model built on a Mixture-of-Experts architecture, utilizing 235 billion total parameters with 22 billion active parameters per forward pass. Designed specifically for deep reasoning, the model employs a specialized thinking-only mode that forces structured, step-by-step analysis before generating a final response. By leveraging 128 experts—with 8 active at any given time—it achieves high efficiency while maintaining the capacity to handle intricate tasks in mathematics, coding, and scientific research. Its native support for a 262,144-token context window allows it to process and synthesize vast amounts of information, making it a robust choice for users requiring high-fidelity logical output.
The model represents a significant evolution in the Qwen3 series, benefiting from extensive pre-training and post-training stages that refine its instruction-following and alignment capabilities. Through iterative scaling of its reasoning depth, the model has reached state-of-the-art performance among open-source alternatives, excelling in academic benchmarks that demand human-level expertise. Its design is optimized for complex agentic workflows and tool usage, where the model must navigate multi-step problems. With its ability to handle long-form generation and nuanced technical queries, it serves as a powerful foundation for developers looking to deploy sophisticated reasoning agents that require both reliability and deep analytical precision.
Quick Info
Powered by- Provider
- Kilo Gateway
- Model key
- qwen/qwen3-235b-a22b-thinking-2507
- Release date
- Jul 25, 2025
- Last updated
- Jul 25, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.23
- Output token cost
- $2.30
Limits
- Output tokens
- 117,964 tokens
- Context window
- 131,072 tokens
Transparent token rates
Compare Qwen: Qwen3 235B A22B Thinking 2507 (retires Oct 9) pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.