StanceBench - Sincerity and Honesty: leaderboard

Metric: AUROC (0-1 scale) of the judge's conversation-level signed stance score in separating the 135 role-prompted positive-pole and negative-pole single-speaker instances of the Sincerity and Honesty dimension (S4), 25 percent sample of the Seamless Interaction Improvised subset; binary pole choice asked in both pole orders, malformed outputs dropped. Source: arxiv.org. Saturation forecast: Around December 2026. 5 models tracked.

Top models

#ModelScore
1Gemini 2.5 Flash (Thinking)0.66
2Qwen2.5-Omni-7B0.64

Interactive version: theaggregate.ai/benchmark?slug=stancebench-sincerity-and-honesty · How It Works · Data refreshed daily, snapshot 2026-09-29.