AbstentionBench - underspecified context - QASPER - Precision: leaderboard

Metric: Abstention Precision (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1Llama-3.1-Tulu-3-8B-DPO29.07
2Llama-3.1-Tulu-3-8B28.79
3Llama-3.1-Tulu-3-70B-DPO26.64
4DeepSeek R1 Distill Llama 70B26.47
5Qwen 2.5 32B Instruct26.28
6Llama 3.3 70B Instruct26.03
7Llama 3.1 8B Instruct25.99
8Llama-3.1-Tulu-3-70B25.43
9Mistral 7B Instruct (v0.3)25.08
10Llama 3.1 405B Instruct25
11GPT-4o23.68
12Llama 3.1 70B Instruct23.59
13Llama-3.1-Tulu-3-8B-SFT23.14
14Gemini 1.5 Pro21.49
15Llama 3.1 70B20.91

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-context-qasper-precision · How It Works · Data refreshed daily, snapshot 2026-09-19.