AbstentionBench - false premise - FalseQA - Precision: leaderboard

Metric: Abstention Precision (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1DeepSeek R1 Distill Llama 70B93.63
2GPT-4o90.59
3Qwen 2.5 32B Instruct89.07
4Llama 3.1 70B Instruct88.67
5Llama 3.3 70B Instruct88.35
6Llama-3.1-Tulu-3-70B88.22
7Mistral 7B Instruct (v0.3)87.62
8Llama-3.1-Tulu-3-8B87.54
9Llama-3.1-Tulu-3-70B-DPO87.44
10Llama 3.1 405B Instruct86.25
11Llama-3.1-Tulu-3-8B-DPO85.8
12O184.99
13Llama-3.1-Tulu-3-8B-SFT84.38
14Llama 3.1 8B Instruct83.46
15Gemini 1.5 Pro83.45

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-false-premise-falseqa-precision · How It Works · Data refreshed daily, snapshot 2026-09-19.