AbstentionBench - underspecified context - WorldSense - F1: leaderboard

Metric: Abstention F1 (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1O193.85
2DeepSeek R1 Distill Llama 70B91.47
3Llama 3.1 70B84.66
4Llama 3.1 8B77.5
5Qwen 2.5 32B Instruct76
6Llama 3.1 405B Instruct75.97
7GPT-4o75.45
8Llama 3.3 70B Instruct73.47
9Llama 3.1 70B Instruct67.09
10Llama-3.1-Tulu-3-70B63.99
11Llama-3.1-Tulu-3-70B-DPO62.79
12Gemini 1.5 Pro60.63
13Mistral 7B Instruct (v0.3)54.88
14Llama-3.1-Tulu-3-8B-DPO51.73
15Llama-3.1-Tulu-3-8B51.23

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-context-worldsense-f1 · How It Works · Data refreshed daily, snapshot 2026-09-19.