Venice AI
See how Qwen 3 235B A22B Instruct 2507 ranks on Sonar's LLM leaderboard across code quality, security, and complexity.
Model details
Qwen 3 235B A22B Instruct 2507 is the dedicated non-thinking variant developed alongside a separate reasoning model from the same 235B foundation. This split design choice reflects a deliberate push to optimize one model for fast, direct response generation without chain-of-thought overhead. The model sits at the frontier of non-reasoning performance, outranking major closed models like GPT-4.1 and Claude Opus 4 on the Artificial Analysis Intelligence Index—a blended benchmark spanning seven areas including general knowledge, reasoning, coding, and STEM. Its architecture is tuned for instruction following, logical reasoning, text comprehension, mathematics, science, and tool usage, making it well-suited for interactive chat, agentic pipelines, and production workflows that demand reliable and immediate generation rather than extended deliberation.
The lineage traces back to the Qwen family, with this 2507 iteration advancing from earlier Qwen3 256B hybrid checkpoints by improving general capabilities and long-tail performance across diverse domains. Developer feedback shaped the final split into thinking and non-thinking tracks, with this version sharpened for latency-sensitive and cost-conscious deployments. Tool calling is a native strength, supporting multi-step agentic tasks that chain calendar, memory, and scheduling operations in parallel. Deployed on Cerebras infrastructure, the model serves with FP8 weights and reaches generation speeds that shatter typical GPU cloud benchmarks, making it a practical choice for high-throughput applications without sacrificing broad competency.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Venice AI
See how Qwen 3 235B A22B Instruct 2507 ranks on Sonar's LLM leaderboard across code quality, security, and complexity.