AbstentionBench - underspecified context - GPQA-Diamond - Precision: leaderboard

Metric: Abstention Precision (%). Source: github.com. 21 models tracked.

Top models

#ModelScore
1Llama 3.1 8B Instruct100
2GPT-4o100
3Llama 3.1 405B Instruct100
4Qwen 2.5 32B Instruct100
5O1100
6Llama-3.1-Tulu-3-8B100
7Llama-3.1-Tulu-3-8B-DPO100
8O1 (Low)100
9Llama-3.1-Tulu-3-70B96.3
10Llama-3.1-Tulu-3-70B-DPO96.3
11Gemini 1.5 Pro95.24
12Llama 3.1 70B Instruct94.74
13Llama 3.3 70B Instruct93.75
14O1 (High)92.31
15Llama 3.1 70B90

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-context-gpqa-diamond-precision · How It Works · Data refreshed daily, snapshot 2026-09-19.