Sulat.com
AI models
Get 10-25% off
Get 10-25% off from Qwen
Alibaba logo

Model details

Qwen3 235B-A22B

Qwen3 235B-A22B is a mixture-of-experts language model built on the Qwen series lineage, designed to balance broad capability with computational efficiency. With 235 billion total parameters but only 22 billion activated per token, it leverages a sparse architecture that maintains deep knowledge capacity without requiring full-model computation at every forward pass. The model spans 94 transformer layers with grouped query attention, using 64 query heads and 4 key-value heads per layer—a configuration that preserves the ability to follow complex reasoning chains while keeping memory and compute manageable. What truly differentiates this model is its seamless switching between thinking mode for extended logical, mathematical, and coding tasks and non-thinking mode for rapid, general-purpose dialogue—all within a single unified framework rather than separate specialized models.

The model was developed through extensive pretraining followed by post-training that prioritized human preference alignment, instruction following, and agentic tool use. This training lineage enables Qwen3 to surpass the earlier QwQ reasoning model and the Qwen2.5 instruction series across mathematics, code generation, and commonsense reasoning benchmarks. Its multilingual foundation supports over 100 languages and dialects with strong instruction-following and translation capabilities across that range. The model excels at agent-based workflows, enabling precise external tool integration in both reasoning and conversational modes, achieving leading open-source performance on complex multi-step tasks. Organizations adopting this model gain an open-weight solution that combines deep reasoning depth with practical versatility for production AI systems, multilingual deployments, and agent orchestration at scale.

Alibabaqwen3-235b-a22bqwen

Quick Info

Powered by
Provider
Alibaba
Model key
qwen3-235b-a22b
Release date
Apr 1, 2025
Last updated
Apr 1, 2025
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.70
Output token cost
$2.80

Limits

Output tokens
16,384 tokens
Context window
131,072 tokens

Transparent token rates

Compare Qwen3 235B-A22B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3 235B-A22B

Alibaba

CoverageBenchmark

Lyceum Technology's hosting page confirms that Qwen3-235B-A22B is a Mixture-of-Experts model developed by Alibaba Cloud's Qwen team with 235 billion total parameters and 22 billion active per forward pass, and identifies the Instruct-2507 checkpoint as optimized for general-purpose text generation, coding, and tool usa The Lyceum page sets Qwen/Qwen3-235B-A22B-Instruct-2507 pricing at $0.20 per million input tokens and $0.60 per million output tokens, billed per token with no base fees, service tiers, or minimum commitments, making rates predictable across bursty and sustained workloads. It frames the offering as enabling EU data-res

Alibaba

Coverage

The Stanford CRFM Foundation Model Transparency Index entry for Alibaba summarizes the Qwen 3 technical report's data acquisition disclosures, noting that the most relevant document is the Qwen 3 technical report. It states Qwen3 models were trained on 36 trillion tokens spanning 119 languages and dialects across domai The same Transparency Index entry outlines Qwen3's pre-training stages, including a reasoning stage (S2) that increases the proportion of STEM, coding, reasoning, and synthetic data, followed by a long-context stage that pre-trains on hundreds of billions of tokens at a sequence length of 32,768 tokens to extend contex

Alibaba

CoverageBenchmark

OpenRouter's listing for Qwen3-235B-A22B-Instruct-2507 confirms it is a multilingual, instruction-tuned mixture-of-experts model with 235B total parameters and 22B active per forward pass, optimized for instruction following, logical reasoning, math, code, and tool usage. The page specifies a native 262K context length The same OpenRouter page provides a developer-focused routing matrix across roughly ten providers, including GMICloud, NovitaAI, Parasail, Alibaba Cloud International, AtlasCloud, StreamLake, Google Vertex (ZDR), Nebius Token Factory, DeepInfra, and Venice, with per-token input/output and cache read rates plus latency,

Videos about Qwen3 235B-A22B

More models around Qwen3 235B-A22B