Qwen3-Omni-30B-A3B (Thinking): benchmark results

Provider: Alibaba. Access: Open.

Unified ELO 1530 ± 1, rank #797 of 2033 rated models, from 58 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
MOV-Bench55.49Accuracy (%; 519 audio-visual multi-hop multiple-choice ques100
MOV-Bench - Intent Reasoning58.75Accuracy (%; 80 intent-reasoning questions; direct inference100
MOV-Bench - Relational Reasoning60.63Accuracy (%; 127 relational-reasoning questions; direct infe100
OmniClean34.93Accuracy (%; unweighted mean over nine audio-visual benchmar100
OmniClean - Daily-Omni42.62Accuracy (%; the 237 of 1,197 Daily-Omni queries left after 100
OmniClean - IntentBench36.42Accuracy (%; the 660 of 2,689 IntentBench queries left after100
OmniClean - UNO-Bench37.55Accuracy (%; the 228 of 1,000 UNO-Bench multiple-choice (UNO100
OmniClean - Video-Holmes46.33Accuracy (%; the 885 of 1,837 Video-Holmes queries left afte100
OmniClean - WorldSense27.7Accuracy (%; the 875 of 3,172 WorldSense queries left after 100
SEA-SpeechBench - Age Recognition (SEA Prompt)36.78Macro-F1 (%; speaker age group (teens, adults, seniors) from100
SEA-SpeechBench - Age Recognition (English Prompt)38.1Macro-F1 (%; speaker age group (teens, adults, seniors) from92.9
SEA-SpeechBench - Timestamped Content Query (0-30 s)1.28WER or CER (%; transcribe only the speech inside a queried t92.9

Interactive version: theaggregate.ai/model?slug=qwen3-omni-30b-a3b-thinking · How It Works · Data refreshed daily, snapshot 2026-09-27.