NovitaAI
A community best-practices guide posted on the official Qwen/Qwen3-Omni-30B-A3B-Thinking HuggingFace repository (Sep 26, 2025) outlines how to evaluate Qwen3-Omni using the EvalScope framework. The guide notes that multimodal inference requires an OpenAI API-compatible serving endpoint, as local transformers inference The discussion provides concrete deployment commands (e.g., a vLLM serve command exposing http://localhost:8801/v1) and a stress-test recipe on an A100 80G GPU with 1024 text tokens plus a 512x512 image as input and 1024 tokens as output. It notes that Qwen3-Omni-30B-A3B-Instruct and Qwen3-30B-A3B-Instruct share the sa