AbstentionBench - subjective - KUQ/Controversial - Precision: leaderboard

Metric: Abstention Precision (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1DeepSeek R1 Distill Llama 70B96.6
2Llama-3.1-Tulu-3-8B-SFT96.15
3Llama-3.1-Tulu-3-8B94.52
4Llama 3.3 70B Instruct93.09
5Llama-3.1-Tulu-3-8B-DPO92.41
6O192.21
7Mistral 7B Instruct (v0.3)91.58
8GPT-4o90.8
9Llama-3.1-Tulu-3-70B-DPO90.6
10Llama 3.1 70B Instruct90.37
11Llama-3.1-Tulu-3-70B89.96
12Llama 3.1 70B88.89
13Llama 3.1 405B Instruct88.52
14Gemini 1.5 Pro86.69
15Llama 3.1 8B86.41

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-subjective-kuq-controversial-precision · How It Works · Data refreshed daily, snapshot 2026-09-19.