AbstentionBench - underspecified context - MMLU History - Recall: leaderboard

Metric: Abstention Recall (%). Source: github.com. 21 models tracked.

Top models

#ModelScore
1Gemini 1.5 Pro58.97
2GPT-4o51.28
3Qwen 2.5 32B Instruct48.72
4O143.59
5Llama-3.1-Tulu-3-70B-DPO41.03
6Llama 3.1 8B Instruct38.46
7Llama 3.1 70B Instruct35.9
8Llama-3.1-Tulu-3-70B35.9
9O1 (High)35.9
10O1 (Low)35.9
11Llama 3.1 405B Instruct33.33
12TinyLlama-1.1B-Chat-v1.023.08
13Llama 3.3 70B Instruct20.51
14Llama-3.1-Tulu-3-8B15.38
15Llama-3.1-Tulu-3-8B-DPO15.38

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-context-mmlu-history-recall · How It Works · Data refreshed daily, snapshot 2026-09-19.