AEQ-Bench - Discrimination: leaderboard

Metric: Discrimination (1-6; how far the responses to one utterance under different contexts (Meld, GigaSpeech), or to one context in different emotional tones (EmoV-DB), are tailored to each variant, rated by a GPT-5-mini judge; mean over the three subsets). Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.

Top models

#ModelScore
1Qwen-Audio-Chat4.86
2Qwen3 Omni 30B A3B Instruct4.71
3SALMONN-7B4.39
4Qwen2.5-Omni-7B4.17
5Qwen2-Audio-7B-Instruct4

Interactive version: theaggregate.ai/benchmark?slug=aeq-bench-discrimination · How It Works · Data refreshed daily, snapshot 2026-09-26.