AbstentionBench - underspecified context - BB/Disambiguate - Precision: leaderboard

Metric: Abstention Precision (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1Llama-3.1-Tulu-3-70B72.73
2Llama-3.1-Tulu-3-70B-DPO63.41
3DeepSeek R1 Distill Llama 70B54.05
4Llama 3.1 8B53.57
5Llama-3.1-Tulu-3-8B53.57
6Llama 3.1 70B52.67
7Llama 3.1 405B Instruct48.53
8Mistral 7B Instruct (v0.3)47.62
9O146.67
10Qwen 2.5 32B Instruct46.53
11Llama-3.1-Tulu-3-8B-DPO46.15
12Llama 3.1 70B Instruct45
13Llama 3.3 70B Instruct43.98
14Gemini 1.5 Pro39.74
15Llama-3.1-Tulu-3-8B-SFT37.5

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-context-bb-disambiguate-precision · How It Works · Data refreshed daily, snapshot 2026-09-19.