AbstentionBench - underspecified context - UMWP - Recall: leaderboard

Metric: Abstention Recall (%). Source: github.com. 19 models tracked.

Top models

#ModelScore
1O176.11
2Llama 3.1 70B Instruct70.63
3Gemini 1.5 Pro68.23
4Qwen 2.5 32B Instruct66.23
5GPT-4o66.17
6Llama 3.1 8B Instruct61.66
7Mistral 7B Instruct (v0.3)60.63
8DeepSeek R1 Distill Llama 70B56.69
9Llama 3.1 70B54.15
10Llama 3.3 70B Instruct53.37
11Llama 3.1 8B50.6
12Llama-3.1-Tulu-3-70B49.83
13Llama-3.1-Tulu-3-70B-DPO49.14
14Llama-3.1-Tulu-3-8B-DPO31.77
15Llama-3.1-Tulu-3-8B31.6

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-context-umwp-recall · How It Works · Data refreshed daily, snapshot 2026-09-19.