DEAF - Background Sound-Semantic Conflict: leaderboard

Metric: Acoustic Robustness Score (%): per level, the harmonic mean of accuracy (share of mismatched clips whose open-ended answer a DeepSeek-R1 judge matches to the acoustic ground truth) and acoustic sensitivity (share of items whose judged answer changes between matched and mismatched audio), averaged over three levels of textual interference (L1 conflicting spoken content, L2 a misleading prompt over neutral content, L3 both); each cell is the mean of three zero-shot runs, on the background sound-semantic conflict set (1,260 clips of speech mixed with DEMAND environment recordings at five signal-to-noise ratios; the model names the acoustic environment, not the one the words imply); higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 7 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Flash8.7#93
2Gemini 2.5 Flash7.5#237

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=deaf-background-sound-semantic-conflict · How It Works · Data refreshed daily, snapshot 2026-10-11.