AbstentionBench - subjective - MoralChoice - Recall: leaderboard

Metric: Abstention Recall (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1Llama 3.1 70B57.35
2Llama 3.1 8B38.38
3Llama-3.1-Tulu-3-70B29.71
4Llama-3.1-Tulu-3-70B-DPO29.26
5O129.12
6Llama 3.1 405B Instruct28.09
7Llama-3.1-Tulu-3-8B-DPO27.79
8Mistral 7B Instruct (v0.3)26.62
9Llama-3.1-Tulu-3-8B26.32
10GPT-4o26.18
11Qwen 2.5 32B Instruct26.18
12Llama-3.1-Tulu-3-8B-SFT25.74
13Llama 3.1 8B Instruct25.29
14Llama 3.3 70B Instruct24.26
15Llama 3.1 70B Instruct24.26

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-subjective-moralchoice-recall · How It Works · Data refreshed daily, snapshot 2026-09-19.