AEQ-Bench - Discrimination: leaderboard
Metric: Discrimination (1-6; how far the responses to one utterance under different contexts (Meld, GigaSpeech), or to one context in different emotional tones (EmoV-DB), are tailored to each variant, rated by a GPT-5-mini judge; mean over the three subsets). Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen-Audio-Chat | 4.86 |
| 2 | Qwen3 Omni 30B A3B Instruct | 4.71 |
| 3 | SALMONN-7B | 4.39 |
| 4 | Qwen2.5-Omni-7B | 4.17 |
| 5 | Qwen2-Audio-7B-Instruct | 4 |
Interactive version: theaggregate.ai/benchmark?slug=aeq-bench-discrimination · How It Works · Data refreshed daily, snapshot 2026-09-26.