AbstentionBench - false premise - FalseQA - Recall: leaderboard

Metric: Abstention Recall (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1Gemini 1.5 Pro69
2O165.94
3Qwen 2.5 32B Instruct64.05
4Llama 3.3 70B Instruct60.7
5GPT-4o60.26
6Llama 3.1 405B Instruct56.62
7Llama 3.1 70B Instruct53.57
8Llama-3.1-Tulu-3-70B-DPO51.67
9Mistral 7B Instruct (v0.3)51.53
10Llama-3.1-Tulu-3-70B51.24
11Llama 3.1 8B Instruct47.74
12Llama-3.1-Tulu-3-8B-SFT43.23
13DeepSeek R1 Distill Llama 70B42.79
14Llama-3.1-Tulu-3-8B-DPO41.34
15Llama-3.1-Tulu-3-8B39.88

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-false-premise-falseqa-recall · How It Works · Data refreshed daily, snapshot 2026-09-19.