AbstentionBench - underspecified context - SQuAD 2.0 - Recall: leaderboard

Metric: Abstention Recall (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1Qwen 2.5 32B Instruct74.99
2Llama 3.1 405B Instruct73.35
3Llama 3.3 70B Instruct69.68
4O169.17
5Llama 3.1 70B Instruct63.47
6Gemini 1.5 Pro62.9
7Mistral 7B Instruct (v0.3)61.26
8Llama-3.1-Tulu-3-70B-DPO59.46
9Llama 3.1 8B Instruct58.05
10GPT-4o57.76
11Llama-3.1-Tulu-3-70B57.54
12Llama 3.1 70B55.11
13Llama 3.1 8B52.17
14DeepSeek R1 Distill Llama 70B48.33
15Llama-3.1-Tulu-3-8B-DPO30.38

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-context-squad-2-0-recall · How It Works · Data refreshed daily, snapshot 2026-09-19.