AbstentionBench - underspecified context - BB/Disambiguate - Recall: leaderboard

Metric: Abstention Recall (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1Llama 3.3 70B Instruct97.33
2Llama 3.1 70B Instruct96
3Llama 3.1 70B92
4GPT-4o89.33
5Qwen 2.5 32B Instruct89.33
6Llama 3.1 405B Instruct88
7O184
8Gemini 1.5 Pro82.67
9Llama 3.1 8B Instruct78.67
10Llama 3.1 8B40
11Llama-3.1-Tulu-3-70B-DPO34.67
12DeepSeek R1 Distill Llama 70B26.67
13Llama-3.1-Tulu-3-70B21.33
14Llama-3.1-Tulu-3-8B20
15Llama-3.1-Tulu-3-8B-DPO16

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-context-bb-disambiguate-recall · How It Works · Data refreshed daily, snapshot 2026-09-19.