AOR-Bench - MB-Score: leaderboard

Metric: MB-Score (%; harmonic mean of the true-refusal rate on 500 harmful speech-only queries and 100 minus the over-refusal rate, pooled over six scenario categories). Source: arxiv.org. Saturation forecast: Around 2028. 12 models tracked.

Top models

#ModelScore
1Qwen2-Audio-7B-Instruct82.8
2Gemini 3 Flash (Preview)55.9
3MiMo-V2-Omni55.68
4Gemini 2.0 Flash Lite54.16
5GPT Audio48.17
6Gemini 2.5 Flash Lite47.97
7GPT Audio Mini39.22

Interactive version: theaggregate.ai/benchmark?slug=aor-bench-mb-score · How It Works · Data refreshed daily, snapshot 2026-09-26.