ParaIntent - Intent Fulfillment: leaderboard

Metric: Intent fulfillment score (%; GLM-5 judge of whether the spoken response addresses the user's intent; over the ParaIntent synthetic test set (14 Chinese spoken-emotional-dialogue intent categories, balanced explicit and implicit samples); single-turn spoken response). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 5 models tracked.

Top models

#ModelScore
1Qwen3-Omni52.27
2MiMo-Audio47.95
3Step-Audio-2-mini41.48
4GLM-4-Voice33.44
5Qwen2.5-Omni32.78

Interactive version: theaggregate.ai/benchmark?slug=paraintent-intent-fulfillment · How It Works · Data refreshed daily, snapshot 2026-09-26.