AbstentionBench - underspecified intent - KUQ/Ambiguous - Recall: leaderboard

Metric: Abstention Recall (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1GPT-4o75.83
2Gemini 1.5 Pro75.5
3Qwen 2.5 32B Instruct69.87
4Llama-3.1-Tulu-3-70B-DPO69.54
5Llama-3.1-Tulu-3-70B68.54
6O166.23
7Llama 3.1 8B Instruct65.56
8Llama 3.3 70B Instruct65.56
9Llama 3.1 405B Instruct65.56
10Llama 3.1 70B Instruct63.25
11Mistral 7B Instruct (v0.3)62.91
12Llama-3.1-Tulu-3-8B-DPO62.25
13Llama-3.1-Tulu-3-8B60.6
14Llama-3.1-Tulu-3-8B-SFT54.64
15Llama 3.1 8B48.68

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-intent-kuq-ambiguous-recall · How It Works · Data refreshed daily, snapshot 2026-09-19.