AbstentionBench - underspecified context - MMLU Math - Recall: leaderboard

Metric: Abstention Recall (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1Llama 3.1 70B Instruct73.68
2GPT-4o71.43
3O1 (Low)71.43
4O169.92
5Llama 3.1 405B Instruct69.17
6O1 (High)68.42
7Llama 3.1 8B Instruct65.41
8Qwen 2.5 32B Instruct64.66
9Llama-3.1-Tulu-3-70B-DPO62.41
10Mistral 7B Instruct (v0.3)61.65
11Gemini 1.5 Pro51.13
12Llama 3.3 70B Instruct48.12
13DeepSeek R1 Distill Llama 70B33.83
14Llama-3.1-Tulu-3-8B-DPO30.83
15Llama-3.1-Tulu-3-8B-SFT27.82

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-context-mmlu-math-recall · How It Works · Data refreshed daily, snapshot 2026-09-19.