AbstentionBench - underspecified context - BBQ - Recall: leaderboard

Metric: Abstention Recall (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1Qwen 2.5 32B Instruct95.05
2Gemini 1.5 Pro92.41
3Mistral 7B Instruct (v0.3)90.78
4GPT-4o90.04
5Llama 3.1 8B Instruct89.93
6Llama 3.1 405B Instruct89.82
7Llama-3.1-Tulu-3-70B88.08
8Llama-3.1-Tulu-3-70B-DPO87.85
9Llama 3.1 70B Instruct84.2
10Llama 3.3 70B Instruct79.75
11Llama-3.1-Tulu-3-8B-DPO65.86
12DeepSeek R1 Distill Llama 70B63.05
13Llama-3.1-Tulu-3-8B62.09
14O154.78
15Llama 3.1 70B46.06

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-context-bbq-recall · How It Works · Data refreshed daily, snapshot 2026-09-19.