AbstentionBench - underspecified context - QASPER - F1: leaderboard

Metric: Abstention F1 (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1Llama-3.1-Tulu-3-8B-DPO44.51
2Llama-3.1-Tulu-3-8B44.31
3Llama-3.1-Tulu-3-70B-DPO41.85
4Qwen 2.5 32B Instruct41.4
5DeepSeek R1 Distill Llama 70B41.03
6Llama 3.3 70B Instruct40.97
7Llama 3.1 8B Instruct40.45
8Llama-3.1-Tulu-3-70B40
9Llama 3.1 405B Instruct39.69
10Mistral 7B Instruct (v0.3)39.57
11GPT-4o38
12Llama 3.1 70B Instruct37.37
13Llama-3.1-Tulu-3-8B-SFT35.33
14Gemini 1.5 Pro35.29
15Llama 3.1 70B33.74

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-context-qasper-f1 · How It Works · Data refreshed daily, snapshot 2026-09-19.