AnyAudio-Judge Bench - Chinese (Dynamic Rubric): leaderboard
Metric: Classification accuracy (%) on the Chinese instruction subset under dynamic rubric prompting (the instruction is decomposed into binary rubric items by Qwen3-30B-A3B-Instruct-2507 and the item-level yes probabilities are aggregated): judge a speech, sound, music or mixed-audio clip against a text instruction as aligned or misaligned, balanced 1:1 positives and negatives (chance 50); accuracy averaged over the seven subsets (Speech-Real, Speech-Gen, Sound-Real, Sound-Gen, Music-Real, Music-Gen, Mix); higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Pro | 78.31 |
| 2 | Qwen3 Omni 30B A3B Instruct | 76.82 |
| 3 | Qwen2.5-Omni-7B | 71.93 |
Interactive version: theaggregate.ai/benchmark?slug=anyaudio-judge-bench-chinese-dynamic-rubric · How It Works · Data refreshed daily, snapshot 2026-09-29.