DEAF - Emotion-Semantic Conflict: leaderboard

Metric: Acoustic Robustness Score (%): per level, the harmonic mean of accuracy (share of mismatched clips whose open-ended answer a DeepSeek-R1 judge matches to the acoustic ground truth) and acoustic sensitivity (share of items whose judged answer changes between matched and mismatched audio), averaged over three levels of textual interference (L1 conflicting spoken content, L2 a misleading prompt over neutral content, L3 both); each cell is the mean of three zero-shot runs, on the emotion-semantic conflict set (1,248 clips of 104 sentences spoken in angry, happy, neutral or sad prosody; the model names the speaker's emotion from the voice, not the words); higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 7 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Flash12.7#237
2Gemini 3 Flash2.6#93

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=deaf-emotion-semantic-conflict · How It Works · Data refreshed daily, snapshot 2026-10-11.