AnyAudio-Judge Bench - Chinese (Holistic): leaderboard
Metric: Classification accuracy (%) on the Chinese instruction subset under holistic prompting (one match or mismatch decision per pair): judge a speech, sound, music or mixed-audio clip against a text instruction as aligned or misaligned, balanced 1:1 positives and negatives (chance 50); accuracy averaged over the seven subsets (Speech-Real, Speech-Gen, Sound-Real, Sound-Gen, Music-Real, Music-Gen, Mix); higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Pro | 80.01 |
| 2 | Qwen3 Omni 30B A3B Instruct | 60.24 |
| 3 | Qwen2.5-Omni-7B | 51.76 |
Interactive version: theaggregate.ai/benchmark?slug=anyaudio-judge-bench-chinese-holistic · How It Works · Data refreshed daily, snapshot 2026-09-29.