DEAF - Speaker Identity-Semantic Conflict: leaderboard
Metric: Acoustic Robustness Score (%): per level, the harmonic mean of accuracy (share of mismatched clips whose open-ended answer a DeepSeek-R1 judge matches to the acoustic ground truth) and acoustic sensitivity (share of items whose judged answer changes between matched and mismatched audio), averaged over three levels of textual interference (L1 conflicting spoken content, L2 a misleading prompt over neutral content, L3 both); each cell is the mean of three zero-shot runs, on the speaker identity-semantic conflict set (248 synthesized clips; the model names the speaker's gender and age from the voice, not from what is said); higher is better. Source: arxiv.org. Saturation forecast: Around May 2027. 7 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Gemini 2.5 Flash | 46.8 | #237 |
| 2 | Gemini 3 Flash | 39.8 | #93 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=deaf-speaker-identity-semantic-conflict · How It Works · Data refreshed daily, snapshot 2026-10-11.