ParaIntent - Emotion Accuracy: leaderboard

Metric: Emotion accuracy (%; MiDashengLM emotion-label agreement between the spoken response and the target emotion; over the ParaIntent synthetic test set (14 Chinese spoken-emotional-dialogue intent categories, balanced explicit and implicit samples); single-turn spoken response). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 5 models tracked.

Top models

#ModelScore
1Step-Audio-2-mini57.88
2Qwen3-Omni51.18
3MiMo-Audio48.39
4Qwen2.5-Omni45.94
5GLM-4-Voice40.24

Interactive version: theaggregate.ai/benchmark?slug=paraintent-emotion-accuracy · How It Works · Data refreshed daily, snapshot 2026-09-26.