AbstentionBench - false premise - CoCoNot/False presumptions - Recall: leaderboard

Metric: Abstention Recall (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1O187.95
2Llama-3.1-Tulu-3-70B80.72
3Llama 3.1 405B Instruct79.52
4Llama-3.1-Tulu-3-70B-DPO79.52
5Llama-3.1-Tulu-3-8B-DPO75.9
6GPT-4o74.7
7Qwen 2.5 32B Instruct74.7
8Llama 3.3 70B Instruct73.49
9Llama-3.1-Tulu-3-8B72.29
10Mistral 7B Instruct (v0.3)71.08
11Llama 3.1 70B Instruct69.88
12Gemini 1.5 Pro69.88
13Llama 3.1 8B Instruct68.67
14DeepSeek R1 Distill Llama 70B66.27
15Llama-3.1-Tulu-3-8B-SFT63.86

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-false-premise-coconot-false-presumptions-recall · How It Works · Data refreshed daily, snapshot 2026-09-19.