AbstentionBench - underspecified context - BB/Disambiguate - F1: leaderboard

Metric: Abstention F1 (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1Llama 3.1 70B66.99
2Llama 3.1 405B Instruct62.56
3Llama 3.1 70B Instruct61.28
4Qwen 2.5 32B Instruct61.19
5Llama 3.3 70B Instruct60.58
6O160
7Gemini 1.5 Pro53.68
8GPT-4o50.76
9Llama 3.1 8B Instruct46.83
10Llama 3.1 8B45.8
11Llama-3.1-Tulu-3-70B-DPO44.83
12DeepSeek R1 Distill Llama 70B35.71
13Llama-3.1-Tulu-3-70B32.99
14Llama-3.1-Tulu-3-8B29.13
15Llama-3.1-Tulu-3-8B-DPO23.76

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-context-bb-disambiguate-f1 · How It Works · Data refreshed daily, snapshot 2026-09-19.