Currently listed through these providers:
Model details
Qwen3 Omni 30B A3B Instruct
Qwen3 Omni 30B A3B Instruct is a natively end-to-end multimodal foundation model designed to process text, images, audio, and video simultaneously. Built on a Mixture of Experts architecture, the model utilizes a Thinker–Talker design with 30 billion total parameters and 3 billion active parameters, which balances high-level reasoning with efficient inference. This design, paired with a multi-codebook approach, minimizes latency to support real-time streaming responses. It is engineered to maintain strong performance across unimodal tasks while achieving state-of-the-art results in audio and video benchmarks, ensuring that its multimodal capabilities do not come at the cost of text or image processing accuracy.
The model benefits from early text-first pretraining combined with mixed multimodal training, which allows it to handle complex, real-world tasks like generating marketing scripts, selecting visual assets, and synthesizing natural-sounding voiceovers. Its training lineage emphasizes broad linguistic support, covering 119 text languages and extensive speech input and output options. By leveraging AuT pretraining for robust general representations, the model excels in multilingual dialogue, code, and mathematics. Its ability to provide immediate, natural turn-taking makes it a practical choice for applications requiring instant customer assistance and coherent, cross-format content creation.
Quick Info
Powered by- Provider
- NovitaAI
- Model key
- qwen/qwen3-omni-30b-a3b-instruct
- Release date
- Sep 24, 2025
- Last updated
- Sep 24, 2025
- Knowledge cutoff
- 2024-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.25
- Output token cost
- $0.97
Limits
- Output tokens
- 16,384 tokens
- Context window
- 65,536 tokens
Latest news about Qwen3 Omni 30B A3B Instruct
No articles yet. Fetch the latest news to show it here.