AbstentionBench - false premise - CoCoNot/False presumptions - Precision: leaderboard

Metric: Abstention Precision (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1DeepSeek R1 Distill Llama 70B98.21
2Llama-3.1-Tulu-3-8B-SFT98.15
3Llama-3.1-Tulu-3-70B97.1
4GPT-4o96.88
5Llama-3.1-Tulu-3-70B-DPO95.65
6Llama 3.3 70B Instruct95.31
7Qwen 2.5 32B Instruct93.94
8Llama-3.1-Tulu-3-8B93.75
9Mistral 7B Instruct (v0.3)93.65
10Llama 3.1 70B Instruct93.55
11Llama-3.1-Tulu-3-8B-DPO91.3
12O191.25
13Gemini 1.5 Pro90.63
14Llama 3.1 405B Instruct88
15Llama 3.1 8B Instruct86.36

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-false-premise-coconot-false-presumptions-precision · How It Works · Data refreshed daily, snapshot 2026-09-19.