Qwen3-235B-A22B-Instruct-2507 is a Mixture-of-Experts large language model with 235 billion total parameters and 22 billion activated parameters per token. The architecture uses 128 experts with 8 activated per forward pass, 94 transformer layers, and grouped query attention with 64 query heads and 4 key-value heads. Designed specifically as a non-thinking model, it does not produce chain-of-thought reasoning blocks and prioritizes speed and direct response quality. Its native context window spans 262,144 tokens and can be extended up to 1,010,000 tokens, making it well-suited for document processing, extended conversations, and applications requiring broad context retention without latency penalties from reasoning loops.
This model represents a post-trained iteration built upon the Qwen3 MoE foundation, with refinements developed following developer feedback to improve instruction following, logical reasoning, text comprehension, mathematics, science, coding, and tool usage capabilities. According to benchmark results across seven evaluations in the Artificial Analysis Intelligence Index, it achieves state-of-the-art results among non-reasoning models, outperforming GPT-4.1, Claude Opus 4, DeepSeek V3, and Kimi K2. The weights are available openly, with FP8 quantized versions enabling efficient deployment. It supports fine-tuning through LoRA-based approaches and runs at over 1,400 tokens per second on specialized hardware, delivering strong price-performance for production applications requiring reliable, high-speed inference at scale.