Currently listed through these providers:
Model details
Qwen3 235B A22B Instruct 2507
Qwen3-235B-A22B-Instruct-2507 is a causal language model built on a mixture-of-experts architecture that utilizes 235 billion total parameters, with 22 billion parameters activated during inference. Designed as a dedicated non-thinking model, it intentionally omits chain-of-thought processing to prioritize speed and directness in its outputs. The model features 94 layers and 128 experts, with 8 experts activated per token, providing a robust foundation for handling diverse tasks ranging from mathematical problem-solving and coding to scientific analysis and tool usage.
Developed through a rigorous process of pre-training and post-training, this model reflects a strategic shift toward separating instruction-tuned models from those designed for complex reasoning. This lineage enables the model to excel in subjective and open-ended tasks, offering improved alignment with user preferences and a more natural generation style. With a native context window of 256,000 tokens, it is well-suited for processing large-scale documents and long-tail knowledge across multiple languages, making it a versatile tool for enterprise applications and advanced research.
Quick Info
Powered by- Provider
- Jiekou.AI
- Model key
- qwen/qwen3-235b-a22b-instruct-2507
- Release date
- Jan 1, 2026
- Last updated
- Jan 1, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.15
- Output token cost
- $0.80
Limits
- Output tokens
- 16,384 tokens
- Context window
- 131,072 tokens