AbstentionBench - underspecified context - QASPER - Recall: leaderboard

Metric: Abstention Recall (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1Gemini 1.5 Pro98.73
2Qwen 2.5 32B Instruct97.47
3Llama-3.1-Tulu-3-70B-DPO97.47
4Llama 3.3 70B Instruct96.2
5GPT-4o96.2
6Llama 3.1 405B Instruct96.2
7O196.2
8Llama-3.1-Tulu-3-8B96.2
9Llama-3.1-Tulu-3-8B-DPO94.94
10Mistral 7B Instruct (v0.3)93.67
11Llama-3.1-Tulu-3-70B93.67
12Llama 3.1 8B Instruct91.14
13DeepSeek R1 Distill Llama 70B91.14
14Llama 3.1 70B Instruct89.87
15Llama 3.1 70B87.34

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-context-qasper-recall · How It Works · Data refreshed daily, snapshot 2026-09-19.