AbstentionBench - false premise - QAQA - Precision: leaderboard

Metric: Abstention Precision (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1DeepSeek R1 Distill Llama 70B92.86
2GPT-4o92.22
3O188.35
4Llama 3.1 70B Instruct85.71
5Llama 3.1 405B Instruct84.62
6Gemini 1.5 Pro83.85
7Llama 3.3 70B Instruct83.69
8Qwen 2.5 32B Instruct83.02
9Llama-3.1-Tulu-3-70B82.72
10Llama-3.1-Tulu-3-70B-DPO81.87
11Mistral 7B Instruct (v0.3)81.41
12Llama 3.1 8B81
13Llama-3.1-Tulu-3-8B77.07
14Llama-3.1-Tulu-3-8B-DPO76.97
15Llama 3.1 70B75.23

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-false-premise-qaqa-precision · How It Works · Data refreshed daily, snapshot 2026-09-19.