AbstentionBench - underspecified context - ALCUNA - Recall: leaderboard

Metric: Abstention Recall (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1Qwen 2.5 32B Instruct84.06
2Llama 3.1 8B Instruct78.23
3Llama-3.1-Tulu-3-70B-DPO77.02
4Llama-3.1-Tulu-3-70B76.96
5Gemini 1.5 Pro74.6
6GPT-4o74.54
7Llama 3.1 70B71.88
8Llama 3.3 70B Instruct70.5
9Llama 3.1 405B Instruct70.44
10Llama 3.1 70B Instruct70.27
11Mistral 7B Instruct (v0.3)67.84
12DeepSeek R1 Distill Llama 70B62.93
13Llama 3.1 8B53.94
14O153.06
15Llama-3.1-Tulu-3-8B-DPO52.42

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-context-alcuna-recall · How It Works · Data refreshed daily, snapshot 2026-09-19.