AbstentionBench - false premise - KUQ/False assumptions - Recall: leaderboard

Metric: Abstention Recall (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1GPT-4o70.3
2Gemini 1.5 Pro67.67
3Llama-3.1-Tulu-3-70B66.17
4Llama-3.1-Tulu-3-70B-DPO65.79
5Llama 3.1 8B Instruct64.66
6Llama 3.1 405B Instruct62.78
7Qwen 2.5 32B Instruct62.41
8Llama 3.3 70B Instruct60.9
9Llama 3.1 70B Instruct57.52
10Llama-3.1-Tulu-3-8B56.77
11Llama-3.1-Tulu-3-8B-DPO55.64
12O154.89
13Mistral 7B Instruct (v0.3)51.5
14DeepSeek R1 Distill Llama 70B47.37
15Llama-3.1-Tulu-3-8B-SFT46.62

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-false-premise-kuq-false-assumptions-recall · How It Works · Data refreshed daily, snapshot 2026-09-19.