Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Crusoe logo

Model details

Qwen3 235B-A22B Instruct 2507

Qwen3-235B-A22B-Instruct-2507 is Alibaba's dedicated non-thinking, non-reasoning Instruct release derived from the Qwen3 235B family, split from its reasoning sibling after the Qwen team responded to developer feedback asking for separate thinking and non-thinking variants. As a Mixture-of-Experts architecture with 235B total parameters and 22B active per token (A22B), the model targets production-grade instruction following rather than chain-of-thought reasoning, making it suitable for general assistants, content generation, structured extraction, and tool-augmented workflows where predictable, fast responses are preferred over deep deliberation.

Third-party evaluations position Qwen3-235B-A22B-Instruct-2507 as a strong contender among non-reasoning models, with Cerebras citing state-of-the-art results in the Artificial Analysis Intelligence Index across general knowledge, reasoning, coding, and STEM benchmarks, outperforming GPT-4.1, Claude Opus 4, DeepSeek V3, and Kimi K2 on that blended measure. The model also brings practical engineering improvements over the prior hybrid Qwen3 235B release, and independent leaderboards show relatively stable quality across conversation depth, though overall composite ranking places it mid-pack compared with newer entrants. Open-weight availability makes the model attractive for self-hosted deployment and fine-tuning in enterprise environments that require on-prem control.

CrusoeQwen/Qwen3-235B-A22B-Instruct-2507qwen

Quick Info

Powered by
Provider
Crusoe
Model key
Qwen/Qwen3-235B-A22B-Instruct-2507
Release date
Jul 21, 2025
Last updated
Jul 21, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.22
Output token cost
$0.80

Limits

Output tokens
16,384 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen3 235B-A22B Instruct 2507 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3 235B-A22B Instruct 2507

Crusoe

Coverage

The ModelScope listing for Qwen3-235B-A22B-Instruct-2507, hosted on Alibaba's ModelScope hub under the apache-2.0 license, mirrors Qwen's official release notes and confirms the same non-thinking-only checkpoint with the 256K long-context enhancement and general improvements in instruction following, reasoning, math, s Architecture details match Qwen's published specs: 235.09B parameters with 22B activated per forward pass, 94 layers, MoE configuration of 128 experts with 8 activated, GQA at 64 Q heads and 4 KV heads, and a 262,144 native context length extendable to about 1,010,000 tokens. The ModelScope card reproduces the same ben

Crusoe

Coverage

Qwen's official Hugging Face model card for Qwen3-235B-A22B-Instruct-2507 documents the updated non-thinking variant of the Qwen3-235B-A22B family, released as an Apache 2.0 licensed checkpoint. The card highlights improvements in instruction following, logical reasoning, text comprehension, mathematics, science, codin The architecture is a Mixture-of-Experts causal language model with 235B total parameters and 22B activated per forward pass, 234B non-embedding parameters across 94 layers, GQA with 64 Q heads and 4 KV heads, 128 experts with 8 activated, and a native 262,144-token context length extendable up to roughly 1,010,000 tok

Videos about Qwen3 235B-A22B Instruct 2507

More models around Qwen3 235B-A22B Instruct 2507