ParaIntent - Emotion Accuracy: leaderboard
Metric: Emotion accuracy (%; MiDashengLM emotion-label agreement between the spoken response and the target emotion; over the ParaIntent synthetic test set (14 Chinese spoken-emotional-dialogue intent categories, balanced explicit and implicit samples); single-turn spoken response). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 5 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Step-Audio-2-mini | 57.88 |
| 2 | Qwen3-Omni | 51.18 |
| 3 | MiMo-Audio | 48.39 |
| 4 | Qwen2.5-Omni | 45.94 |
| 5 | GLM-4-Voice | 40.24 |
Interactive version: theaggregate.ai/benchmark?slug=paraintent-emotion-accuracy · How It Works · Data refreshed daily, snapshot 2026-09-26.