AbstentionBench - underspecified context - GPQA-Diamond - F1: leaderboard

Metric: Abstention F1 (%). Source: github.com. 21 models tracked.

Top models

#ModelScore
1O187.32
2GPT-4o85.71
3Llama 3.1 405B Instruct78.79
4Qwen 2.5 32B Instruct78.79
5Llama-3.1-Tulu-3-70B77.61
6Llama-3.1-Tulu-3-70B-DPO77.61
7O1 (Low)76.92
8O1 (High)72.73
9Llama 3.1 8B Instruct68.85
10Gemini 1.5 Pro65.57
11Mistral 7B Instruct (v0.3)64.62
12Llama 3.1 70B Instruct61.02
13Llama 3.3 70B Instruct53.57
14Llama-3.1-Tulu-3-8B-SFT44.83
15DeepSeek R1 Distill Llama 70B38.46

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-context-gpqa-diamond-f1 · How It Works · Data refreshed daily, snapshot 2026-09-19.