BenchPreS: leaderboard

Metric: Appropriate Application Rate minus Misapplication Rate, in percentage points (-100 to 100): the share of preferences that should be applied and are, minus the share that should be suppressed but are applied, over 1,950 attribute-level instances (DeepSeek-R1 judge, exact match for nicknames; 3 samples at temperature 1.0); higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 10 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5.2 (Thinking)46.38#105 (GPT-5.2)
2Claude Sonnet 4.5 (Thinking)34.94#138 (Claude Sonnet 4.5)
3DeepSeek V3.2 (Thinking)26.48#198 (DeepSeek V3.2)
4K-EXAONE (Thinking)25.82#361 (K-EXAONE)
5Qwen 3 32B (Non-reasoning)13.23#424 (Qwen 3 32B)
6GPT-OSS-120B13.12#330
7Mistral 7B Instruct (v0.3)11.28#1255
8Llama 3.3 70B Instruct10.74#520
9Gemini 3 Pro2.21#77
10Qwen 3 235B A22B 2507 (Thinking)-1.77#253 (Qwen 3 235B A22B 2507)

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=benchpres · How It Works · Data refreshed daily, snapshot 2026-10-11.