StanceBench - Cognitive Attentiveness: leaderboard

Metric: AUROC (0-1 scale) of the judge's conversation-level signed stance score in separating the 69 role-prompted positive-pole and negative-pole single-speaker instances of the Cognitive Attentiveness dimension (S5), 25 percent sample of the Seamless Interaction Improvised subset; binary pole choice asked in both pole orders, malformed outputs dropped. Source: arxiv.org. Saturation forecast: Around December 2026. 5 models tracked.

Top models

#ModelScore
1Qwen2.5-Omni-7B0.8
2Gemini 2.5 Flash (Thinking)0.76

Interactive version: theaggregate.ai/benchmark?slug=stancebench-cognitive-attentiveness · How It Works · Data refreshed daily, snapshot 2026-09-29.