AbstentionBench - underspecified context - SQuAD 2.0 - F1: leaderboard

Metric: Abstention F1 (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1Qwen 2.5 32B Instruct84.13
2Llama 3.1 405B Instruct83.56
3Llama 3.3 70B Instruct81.24
4O180.86
5Llama 3.1 70B Instruct76.91
6Gemini 1.5 Pro76.77
7Mistral 7B Instruct (v0.3)74.54
8Llama-3.1-Tulu-3-70B-DPO73.56
9GPT-4o72.79
10Llama 3.1 8B Instruct72.52
11Llama-3.1-Tulu-3-70B72.22
12Llama 3.1 70B70.9
13Llama 3.1 8B68.04
14DeepSeek R1 Distill Llama 70B64.82
15Llama-3.1-Tulu-3-8B-DPO46.36

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-context-squad-2-0-f1 · How It Works · Data refreshed daily, snapshot 2026-09-19.