AbstentionBench - underspecified context - UMWP - F1: leaderboard

Metric: Abstention F1 (%). Source: github.com. 19 models tracked.

Top models

#ModelScore
1O186.38
2Llama 3.1 70B Instruct82.65
3Gemini 1.5 Pro80.68
4GPT-4o79.56
5Qwen 2.5 32B Instruct79.41
6Llama 3.1 8B Instruct76.09
7Mistral 7B Instruct (v0.3)73.43
8DeepSeek R1 Distill Llama 70B72.25
9Llama 3.3 70B Instruct69.47
10Llama 3.1 70B68.2
11Llama-3.1-Tulu-3-70B66.46
12Llama-3.1-Tulu-3-70B-DPO65.82
13Llama 3.1 8B64.46
14Llama-3.1-Tulu-3-8B-DPO48.22
15Llama-3.1-Tulu-3-8B48.02

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-context-umwp-f1 · How It Works · Data refreshed daily, snapshot 2026-09-19.