Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
NovitaAI logo

Model details

Qwen3 Omni 30B A3B Thinking

Qwen3 Omni 30B A3B Thinking serves as the dedicated reasoning engine within the broader Qwen3-Omni framework. Built on a specialized Thinker-Talker architecture, this model is engineered to act as the cognitive core of the system, unifying diverse inputs—including text, audio, images, and video—into a shared representational space. By utilizing a MoE-based design and advanced chain-of-thought reasoning, it is specifically optimized to solve intricate problems that require cross-modal understanding, such as interpreting complex scientific data or mathematical challenges presented in visual formats.

The model benefits from a robust training lineage that emphasizes early text-first pretraining combined with mixed multimodal training to ensure high performance across all supported data types. This approach allows the model to maintain strong general representations while achieving state-of-the-art results on numerous audio and video benchmarks. Designed for high-stakes analytical environments, the model provides deep, step-by-step reasoning capabilities that make it a powerful tool for research, document drafting, and complex problem-solving where precise, text-based output is required.

NovitaAIqwen/qwen3-omni-30b-a3b-thinking

Quick Info

Powered by
Provider
NovitaAI
Model key
qwen/qwen3-omni-30b-a3b-thinking
Release date
Sep 24, 2025
Last updated
Sep 24, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.25
Output token cost
$0.97

Limits

Output tokens
16,384 tokens
Context window
65,536 tokens

Latest news about Qwen3 Omni 30B A3B Thinking

NovitaAI

CoverageDiscourse

A community best-practices guide posted on the official Qwen/Qwen3-Omni-30B-A3B-Thinking HuggingFace repository (Sep 26, 2025) outlines how to evaluate Qwen3-Omni using the EvalScope framework. The guide notes that multimodal inference requires an OpenAI API-compatible serving endpoint, as local transformers inference The discussion provides concrete deployment commands (e.g., a vLLM serve command exposing http://localhost:8801/v1) and a stress-test recipe on an A100 80G GPU with 1024 text tokens plus a 512x512 image as input and 1024 tokens as output. It notes that Qwen3-Omni-30B-A3B-Instruct and Qwen3-30B-A3B-Instruct share the sa

NovitaAI

Coverage

The Qwen3-Omni Technical Report (arXiv:2509.17765, submitted Sep 22, 2025) introduces Qwen3-Omni as a single multimodal model from Alibaba's Qwen team that maintains state-of-the-art performance across text, image, audio, and video without degradation relative to single-modal counterparts. The paper explicitly names Qw Qwen3-Omni adopts a Thinker-Talker Mixture-of-Experts architecture that unifies perception and generation across modalities. Across 36 audio and audio-visual benchmarks, Qwen3-Omni achieves open-source SOTA on 32 and overall SOTA on 22, outperforming closed-source models including Gemini-2.5-Pro, Seed-ASR, and GPT-4o-T

Videos about Qwen3 Omni 30B A3B Thinking