AbstentionBench - subjective - CoCoNot/Humanizing - Recall: leaderboard

Metric: Abstention Recall (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1Llama-3.1-Tulu-3-8B-SFT97.56
2O196.34
3Llama-3.1-Tulu-3-8B-DPO93.9
4Llama-3.1-Tulu-3-70B-DPO93.9
5Llama 3.1 8B Instruct92.68
6Llama 3.3 70B Instruct92.68
7Llama 3.1 70B Instruct91.46
8Llama-3.1-Tulu-3-70B91.46
9Llama-3.1-Tulu-3-8B90.24
10GPT-4o89.02
11Llama 3.1 405B Instruct89.02
12Gemini 1.5 Pro89.02
13Mistral 7B Instruct (v0.3)85.37
14Qwen 2.5 32B Instruct82.93
15DeepSeek R1 Distill Llama 70B63.41

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-subjective-coconot-humanizing-recall · How It Works · Data refreshed daily, snapshot 2026-09-19.