Model details
qwen3-235b-a22b-instruct-2507
Qwen3-235B-A22B-Instruct-2507 is the July 2025 refresh of Qwen's flagship Mixture-of-Experts instruct model, positioned as the non-thinking-mode sibling to the original Qwen3-235B-A22B release. It is built as a causal language model with 235 billion total parameters, 22 billion activated per inference, 234 billion non-embedding parameters, and 94 layers, paired with grouped-query attention using 64 query heads. The post-training pass emphasized instruction following, logical reasoning, mathematics, science, coding, tool usage, broader long-tail knowledge across multiple languages, and more helpful behavior on subjective and open-ended prompts, while also extending the model's native 256K context handling.
In practical terms, the model is aimed at teams that need a single large instruct model for diverse workloads: long-document analysis, multilingual Q&A and writing, code generation with tool assistance, and reasoning-heavy chat where responses should be direct rather than wrapped in chain-of-thought blocks. The FP8-quantized NIM container published on the NVIDIA NGC catalog suggests an inference path tuned for efficient deployment on NVIDIA hardware, complementing the original Qwen distribution. It fits well as a general-purpose assistant backbone for enterprise applications that value long-context comprehension and broad language coverage without switching between specialized models.
Quick Info
Powered by- Provider
- 302.AI
- Model key
- qwen3-235b-a22b-instruct-2507
- Release date
- Jul 30, 2025
- Last updated
- Jul 30, 2025
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.29
- Output token cost
- $1.143
Limits
- Output tokens
- 65,536 tokens
- Context window
- 128,000 tokens
Latest news about qwen3-235b-a22b-instruct-2507
No articles yet. Fetch the latest news to show it here.