AbstentionBench - subjective - MoralChoice - Precision: leaderboard

Metric: Abstention Precision (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1DeepSeek R1 Distill Llama 70B100
2Llama-3.1-Tulu-3-70B98.54
3Llama-3.1-Tulu-3-70B-DPO98.51
4O198.51
5Llama 3.1 405B Instruct98.45
6GPT-4o98.34
7Llama 3.3 70B Instruct98.21
8Llama 3.1 70B Instruct98.21
9Llama-3.1-Tulu-3-8B-DPO97.93
10Llama-3.1-Tulu-3-8B97.81
11Qwen 2.5 32B Instruct97.8
12Llama-3.1-Tulu-3-8B-SFT97.77
13Llama 3.1 8B Instruct97.73
14Gemini 1.5 Pro97.63
15Mistral 7B Instruct (v0.3)96.79

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-subjective-moralchoice-precision · How It Works · Data refreshed daily, snapshot 2026-09-19.