Alibaba
Alibaba Qwen Team Releases Qwen3-ASR: A Speech Recognition Service Built Upon Qwen3-Omni Achieving Robust Speech Recognition Performance
Model details
Qwen3-ASR Flash is a dedicated speech recognition service in the Qwen3 family, built on top of the Qwen3-Omni multimodal foundation and trained on a large corpus of multimodal data that includes tens of millions of hours of ASR-specific audio. This lineage gives it broad acoustic robustness while keeping it focused on the single task of converting speech into accurate text, rather than serving as a general-purpose conversational model.
The model is designed for practical real-time and offline transcription workloads, including live captioning, meeting and call transcription, and voice-driven command interfaces, and it handles noisy or acoustically complex environments without losing accuracy. A distinguishing capability is its support for eleven languages plus multiple accents, combined with a flexible contextual biasing feature that lets users supply background text in any format to steer the output toward custom vocabulary or domain terms. It also extends beyond standard dictation to recognize sung lyrics against background music and reliably rejects non-speech segments such as silence and background noise.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Alibaba
Alibaba Qwen Team Releases Qwen3-ASR: A Speech Recognition Service Built Upon Qwen3-Omni Achieving Robust Speech Recognition Performance