AbstentionBench - false premise - QAQA - Recall: leaderboard

Metric: Abstention Recall (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1O163.86
2Gemini 1.5 Pro56.49
3Llama-3.1-Tulu-3-70B55.44
4GPT-4o54.04
5Llama-3.1-Tulu-3-70B-DPO52.28
6Llama 3.1 405B Instruct50.18
7Llama-3.1-Tulu-3-8B-DPO48.07
8Qwen 2.5 32B Instruct46.32
9Mistral 7B Instruct (v0.3)44.56
10Llama-3.1-Tulu-3-8B42.46
11Llama 3.1 8B Instruct41.4
12Llama 3.3 70B Instruct41.4
13Llama 3.1 70B Instruct40
14Llama-3.1-Tulu-3-8B-SFT36.14
15DeepSeek R1 Distill Llama 70B31.93

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-false-premise-qaqa-recall · How It Works · Data refreshed daily, snapshot 2026-09-19.