Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
DevPass (LLM Gateway) logo

Model details

Qwen3 235B A22B Instruct (2507)

Qwen3-235B-A22B-Instruct-2507 is a dense mixture-of-experts large language model built by Qwen Alibaba, designed as the non-thinking counterpart to the broader Qwen3 family. With 235 billion total parameters and 22 billion activated per inference, the architecture prioritizes efficiency without sacrificing capability. The model natively handles up to 262,144 tokens, making it well suited for long-document analysis, multi-turn conversations, and complex tasks requiring extended context. Its development emerged from explicit developer feedback, with the team creating separate thinking and non-thinking versions to serve different use cases, and the 2507 release represents an evolution over the earlier Qwen3 256B hybrid model with meaningful gains in instruction following, reasoning, mathematics, science, coding, and multilingual understanding.

Benchmarks paint a compelling picture: the model achieves 83.0 on MMLU-Pro, 77.5 on GPQA, and a standout 70.3 on the AIME25 reasoning evaluation, ranking 3rd on ZebraLogic with a 0.95 accuracy score. In the Artificial Analysis Intelligence Index, it outperforms GPT-4.1, Claude Opus 4, DeepSeek V3, and Kimi K2, landing at the top among non-reasoning models. Arena-Hard v2 scores of 79.2 and WritingBench scores of 85.2 further confirm its strength in open-ended and subjective tasks, while tool-calling benchmarks score 96, reflecting the model's robustness for agentic workflows. FP8 quantized weights are available through NVIDIA NIM containers, enabling efficient GPU-accelerated deployment, and inference speeds exceeding 1,400 tokens per second have been demonstrated on specialized hardware, making this model practical for high-throughput production environments.

DevPass (LLM Gateway)qwen3-235b-a22b-instruct-2507qwen

Quick Info

Powered by
Provider
DevPass (LLM Gateway)
Model key
qwen3-235b-a22b-instruct-2507
Release date
Jul 8, 2025
Last updated
Jul 8, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.09
Output token cost
$0.58

Limits

Output tokens
8,192 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen3 235B A22B Instruct (2507) pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3 235B A22B Instruct (2507)

DevPass (LLM Gateway)

CoverageBenchmark

Traictory's model directory entry lists Qwen3-235B-A22B-Instruct-2507 with a release date of July 21, 2025, 235 billion parameters, and an average benchmark score of 72.1%. The model is described as an updated instruction version with substantial improvements in instruction following, logical reasoning, text comprehens Self-reported benchmarks include GPQA accuracy of 77.5%, Aider-Polyglot at 57.3%, and AIME25 for mathematical reasoning evaluation. The directory lists pricing at $0.15 input / $0.80 output per 1M tokens with a 131.1K max input context and 16.4K max output tokens, and supported features including function calling, stru

DevPass (LLM Gateway)

CoverageBenchmark

Requesty aggregates three inference endpoints serving Qwen3-235B-A22B-Instruct-2507: DeepInfra at $0.07 input / $0.10 output per 1M tokens with 262K context, Parasail at $0.15 / $0.85 with cache read at $0.15, and Nebius AI (EU region) at $0.20 / $0.60 with 128K context. The router uses a single OpenAI-compatible base Benchmark scores sourced from Artificial Analysis show a Coding Index of 22.1%, GPQA Diamond reasoning of 79.0%, and Intelligence Index of 19.9% for the model. The page highlights Qwen3-235B-A22B-Instruct-2507 as an open-weights model from Alibaba with significant improvements over the prior Qwen3-235B-A22B non-thinkin

DevPass (LLM Gateway)

Coverage

DeepInfra hosts Qwen3-235B-A22B-Instruct-2507 as a first-party inference endpoint, confirming the model's availability with an fp8 quantization and a native 262,144-token context window. The page lists per-1M-token pricing of $0.09 input and $0.55 output, with Priority tier (1.5×) and Flex tier (0.8×) variants for work The model is described as the updated non-thinking variant of Qwen3-235B-A22B, featuring significant improvements in instruction following, logical reasoning, text comprehension, mathematics, science, coding, and tool usage. The page exposes adjustable inference parameters including max new tokens (up to ~65K or model'

Videos about Qwen3 235B A22B Instruct (2507)

More models around Qwen3 235B A22B Instruct (2507)