PERSPECTRA - Opinion Matching Stance Accuracy: leaderboard

Metric: Stance accuracy (%): the chosen candidate opinion is on the correct side (pro or con) of the debate, for the opinion-matching task (PERSPECTRA: Kialo debate opinions (100 topics, 762 pro and con opinions) each expanded into five Reddit-style arguments by GPT-4o; 500 sampled evaluation items per task, gold labels derived from the Kialo opinion identifiers and stances); higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 12 models tracked.

Top models

#ModelScoreOverall rank
1DeepSeek R1 Distill Qwen 32B95#640
2Qwen 3 8B92.4#667
3Qwen 2.5 7B Instruct91.2#846
4DeepSeek-R1-Distill-Qwen-7B89.4#1400
5Qwen 3 32B86#424
6DeepSeek R1 Distill Llama 8B84.8#1282
7GPT-4o Mini81.6#588
8GPT-4o81.2#333
9QwQ-32B80.2#410
10Falcon3-7B-Instruct77.4#1284

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=perspectra-opinion-matching-stance-accuracy · How It Works · Data refreshed daily, snapshot 2026-10-11.