AbstentionBench - underspecified context - SQuAD 2.0 - Precision: leaderboard

Metric: Abstention Precision (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1Llama 3.1 70B99.39
2Gemini 1.5 Pro98.5
3DeepSeek R1 Distill Llama 70B98.39
4GPT-4o98.37
5Llama-3.1-Tulu-3-8B-DPO97.82
6Llama 3.1 8B97.78
7Llama-3.1-Tulu-3-8B97.74
8Llama 3.1 70B Instruct97.57
9Llama 3.3 70B Instruct97.4
10O197.3
11Llama-3.1-Tulu-3-8B-SFT97.28
12Llama 3.1 405B Instruct97.09
13Llama-3.1-Tulu-3-70B96.96
14Llama 3.1 8B Instruct96.62
15Llama-3.1-Tulu-3-70B-DPO96.43

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-context-squad-2-0-precision · How It Works · Data refreshed daily, snapshot 2026-09-19.