Kilo Gateway
Qwen3 30B A3B Instruct 2507 pricing: $0.09/M input, $0.30/M output. Compare with 10 similar models, see benchmarks, and find the cheapest provider.
Model details
Qwen3 30B A3B sits within Alibaba's Qwen3 family of open-weight language models, intended for general text-based reasoning and generation workloads that benefit from a mixture-of-experts design. Rather than activating its full parameter count at every step, the model uses selective expert routing so that only a subset of its weights is engaged per token, a pattern the broader Qwen3 line adopted to balance capacity against compute cost. The "A3B" designation reflects this design goal: a sizable total parameter budget paired with a much smaller per-token activation footprint, positioning the model as a mid-scale option that aims to deliver stronger reasoning than comparably sized dense models without the inference cost of a full 30B dense run.
In practice, the model is aimed at developers and teams who want open-weight flexibility for chat, instruction following, and tool-assisted workflows, while still expecting reasonable throughput from a moderately sized deployment. Its place in the Qwen3 lineup makes it a sensible choice for experimentation, fine-tuning, and integration into pipelines where a sparse model with a smaller active parameter count is preferable to a dense counterpart. Users evaluating it should weigh its MoE routing behavior and mid-range scale against their latency, memory, and quality targets rather than treating it as a drop-in replacement for either a small dense model or the largest Qwen3 variants.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Kilo Gateway
Qwen3 30B A3B Instruct 2507 pricing: $0.09/M input, $0.30/M output. Compare with 10 similar models, see benchmarks, and find the cheapest provider.