AnyAudio-Judge Bench - Chinese (Holistic): leaderboard

Metric: Classification accuracy (%) on the Chinese instruction subset under holistic prompting (one match or mismatch decision per pair): judge a speech, sound, music or mixed-audio clip against a text instruction as aligned or misaligned, balanced 1:1 positives and negatives (chance 50); accuracy averaged over the seven subsets (Speech-Real, Speech-Gen, Sound-Real, Sound-Gen, Music-Real, Music-Gen, Mix); higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 7 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro80.01
2Qwen3 Omni 30B A3B Instruct60.24
3Qwen2.5-Omni-7B51.76

Interactive version: theaggregate.ai/benchmark?slug=anyaudio-judge-bench-chinese-holistic · How It Works · Data refreshed daily, snapshot 2026-09-29.