Clarifai
Qwen3-30B-A3B-Instruct-2507. New model update from Qwen, improving on their previous Qwen3-30B-A3B release from late April. In their tweet ...
Model details
Qwen3 30B A3B Instruct 2507 is a refresh of the Qwen family's earlier 30B A3B instruct release from late April, positioned as an updated conversational and instruction-following model in the Qwen3 line. Independent commentary describes it as a new model update from Qwen that improves on the prior Qwen3-30B-A3B variant, and community activity confirms its open-weights availability through derivative artifacts such as an AWQ INT4 quantized version (group size 32, calibrated on nvidia/Llama-Nemotron-Post-Training-Dataset via llm-compressor), which compresses the base model from roughly 57 GB to about 17 GB. This combination of public weights and active community quantization work makes the model attractive to teams that want to self-host or fine-tune without paying proprietary licensing costs.
In practical use, Qwen3 30B A3B Instruct 2507 is built around text input and output with a very large 262,144-token context window and matching 262,144-token maximum output, which suits long-document reasoning, multi-turn agent sessions, and retrieval-heavy workflows. It exposes tool calling and temperature control through an OpenAI-compatible chat completions endpoint, allowing structured outputs and programmatic orchestration alongside generation. The model is therefore a strong fit for applications that need long-context instruct behavior, agent or function-calling integration, and the flexibility of an open-weights base, while teams wanting lighter deployment footprints can rely on the community-quantized AWQ variant for memory-constrained environments.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Clarifai
Qwen3-30B-A3B-Instruct-2507. New model update from Qwen, improving on their previous Qwen3-30B-A3B release from late April. In their tweet ...
Clarifai
A community-uploaded Ollama mirror by user "alibayram" republishes the Qwen3-30B-A3B-Instruct-2507 model card, confirming core architecture details: 30.5B total parameters with 3.3B activated across 48 layers, 128 experts with 8 activated, GQA configuration of 32 query and 4 KV attention heads, and a native 262,144-tok The page also reproduces Qwen's own description of the 2507 update's enhancements: gains in instruction following, logical reasoning, text comprehension, mathematics, science, coding and tool usage; expanded long-tail multilingual knowledge; better alignment on subjective and open-ended tasks; and enhanced 256K long-co
Clarifai
The Hugging Face model card for Qwen/Qwen3-30B-A3B-Instruct-2507 is the canonical first-party artifact for the exact model subject. It documents the architecture as a causal language model with 30.5B total parameters and 3.3B activated across 48 layers, using 128 experts with 8 active per token, GQA with 32 Q heads and The same model card reports a comparative benchmark table pitting Qwen3-30B-A3B-Instruct-2507 against DeepSeek-V3-0324, GPT-4o-0327, Gemini-2.5-Flash Non-Thinking, Qwen3-235B-A22B Non-Thinking, and Qwen3-30B-A3B Non-Thinking. Highlighted scores include MMLU-Pro 78.4, GPQA 70.4, AIME25 61.3, HMMT25 43.0, ZebraLogic 90.0