AbstentionBench - underspecified context - WorldSense - Recall: leaderboard

Metric: Abstention Recall (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1O198.96
2DeepSeek R1 Distill Llama 70B97.45
3Llama 3.1 70B94.21
4Llama 3.1 8B90.51
5Llama 3.3 70B Instruct84.14
6Llama-3.1-Tulu-3-70B83.8
7Mistral 7B Instruct (v0.3)83.33
8Llama-3.1-Tulu-3-70B-DPO81.25
9Llama-3.1-Tulu-3-8B-DPO76.27
10Qwen 2.5 32B Instruct74.77
11Llama-3.1-Tulu-3-8B72.22
12GPT-4o67.94
13Llama 3.1 70B Instruct67.48
14Llama 3.1 405B Instruct67.13
15Llama-3.1-Tulu-3-8B-SFT62.85

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-context-worldsense-recall · How It Works · Data refreshed daily, snapshot 2026-09-19.