AbstentionBench - underspecified intent - KUQ/Ambiguous - Precision: leaderboard

Metric: Abstention Precision (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1DeepSeek R1 Distill Llama 70B82.63
2Llama 3.1 70B79.78
3Llama-3.1-Tulu-3-8B-SFT79.33
4Llama-3.1-Tulu-3-8B79.22
5O179.05
6GPT-4o77.63
7Llama-3.1-Tulu-3-8B-DPO76.73
8Llama 3.1 8B76.56
9Gemini 1.5 Pro74.27
10Llama 3.3 70B Instruct73.88
11Llama-3.1-Tulu-3-70B70.89
12Llama-3.1-Tulu-3-70B-DPO70.71
13Mistral 7B Instruct (v0.3)70.63
14Qwen 2.5 32B Instruct69.18
15Llama 3.1 70B Instruct67.49

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-intent-kuq-ambiguous-precision · How It Works · Data refreshed daily, snapshot 2026-09-19.