Alibaba (China)
Automatically distinguish 11 languages and filter noise. Try it for free now!
Model details
Qwen3-ASR-Flash is Alibaba's automatic speech recognition service within the Qwen family, designed to turn spoken audio into text across a wide range of languages and acoustic conditions. According to third-party coverage, the model is built on the Qwen3-Omni foundation and trained on tens of millions of hours of multimodal speech data, giving it a lineage tied to Alibaba's broader multimodal research effort. Its practical focus is on real-world transcription rather than clean studio audio: it targets noisy and far-field recordings, can transcribe lyrics layered over background music, and is able to filter silence and non-speech segments so downstream text stays usable.
The model is positioned for global, mixed-language use, covering eleven languages that include English, Arabic, French, German, Spanish, Italian, Portuguese, Russian, Japanese, Korean, and Chinese with regional dialects such as Cantonese, Sichuanese, Minnan, and Wu. Automatic language detection removes the need for manual configuration when audio switches between languages, which makes it well suited to live captioning, customer call transcription, and multilingual media workflows. It also accepts arbitrary context text covering names, jargon, and domain-specific terminology to bias recognition toward specialized vocabulary, and independent reviewers highlight its low-latency streaming behavior, pointing to a design that balances speed with accuracy for production transcription pipelines.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Alibaba (China)
Automatically distinguish 11 languages and filter noise. Try it for free now!