AbstentionBench - underspecified context - MMLU History - Precision: leaderboard

Metric: Abstention Precision (%). Source: github.com. 21 models tracked.

Top models

#ModelScore
1GPT-4o100
2Llama 3.1 70B100
3O1100
4Llama-3.1-Tulu-3-8B-SFT100
5Llama-3.1-Tulu-3-8B100
6Llama-3.1-Tulu-3-70B100
7Llama-3.1-Tulu-3-8B-DPO100
8Llama-3.1-Tulu-3-70B-DPO100
9O1 (High)100
10O1 (Low)100
11Llama 3.1 70B Instruct93.33
12Gemini 1.5 Pro92
13Llama 3.3 70B Instruct88.89
14Llama 3.1 8B Instruct88.24
15Qwen 2.5 32B Instruct86.36

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-context-mmlu-history-precision · How It Works · Data refreshed daily, snapshot 2026-09-19.