StanceBench - Assertiveness: leaderboard

Metric: AUROC (0-1 scale) of the judge's conversation-level signed stance score in separating the 91 role-prompted positive-pole and negative-pole single-speaker instances of the Assertiveness dimension (S3), 25 percent sample of the Seamless Interaction Improvised subset; binary pole choice asked in both pole orders, malformed outputs dropped. Source: arxiv.org. Saturation forecast: Around December 2026. 5 models tracked.

Top models

#ModelScore
1Gemini 2.5 Flash (Thinking)0.87
2Qwen2.5-Omni-7B0.8

Interactive version: theaggregate.ai/benchmark?slug=stancebench-assertiveness · How It Works · Data refreshed daily, snapshot 2026-09-29.