Qwen3-235B-A22B-Instruct-2507 is a dense mixture-of-experts large language model built by Qwen Alibaba, designed as the non-thinking counterpart to the broader Qwen3 family. With 235 billion total parameters and 22 billion activated per inference, the architecture prioritizes efficiency without sacrificing capability. The model natively handles up to 262,144 tokens, making it well suited for long-document analysis, multi-turn conversations, and complex tasks requiring extended context. Its development emerged from explicit developer feedback, with the team creating separate thinking and non-thinking versions to serve different use cases, and the 2507 release represents an evolution over the earlier Qwen3 256B hybrid model with meaningful gains in instruction following, reasoning, mathematics, science, coding, and multilingual understanding.
Benchmarks paint a compelling picture: the model achieves 83.0 on MMLU-Pro, 77.5 on GPQA, and a standout 70.3 on the AIME25 reasoning evaluation, ranking 3rd on ZebraLogic with a 0.95 accuracy score. In the Artificial Analysis Intelligence Index, it outperforms GPT-4.1, Claude Opus 4, DeepSeek V3, and Kimi K2, landing at the top among non-reasoning models. Arena-Hard v2 scores of 79.2 and WritingBench scores of 85.2 further confirm its strength in open-ended and subjective tasks, while tool-calling benchmarks score 96, reflecting the model's robustness for agentic workflows. FP8 quantized weights are available through NVIDIA NIM containers, enabling efficient GPU-accelerated deployment, and inference speeds exceeding 1,400 tokens per second have been demonstrated on specialized hardware, making this model practical for high-throughput production environments.