Hear2Act (Text, Explicit Feedback): leaderboard

Metric: Optimal-solution rate (%; share of dialogues ending with the first-tier option, using the final recommendation if the turn budget runs out; 480 persona-grounded consumer-service scenarios with a hidden user concern and verifiable outcomes; explicit lexical feedback: the concern is stated in the user's words; text LLM reading the dialogue transcript only, two runs per scenario). Source: arxiv.org. Saturation forecast: Around December 2026. 5 models tracked.

Top models

#ModelScore
1Claude Opus 4.674.9
2Kimi K2.569.4
3GLM-567.3
4Qwen 3 32B64.3
5DeepSeek V3.254.8

Interactive version: theaggregate.ai/benchmark?slug=hear2act-text-explicit-feedback · How It Works · Data refreshed daily, snapshot 2026-09-29.