AbstentionBench - underspecified intent - KUQ/Ambiguous - F1: leaderboard

Metric: Abstention F1 (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1GPT-4o76.72
2Gemini 1.5 Pro74.88
3O172.07
4Llama-3.1-Tulu-3-70B-DPO70.12
5Llama-3.1-Tulu-3-70B69.7
6Qwen 2.5 32B Instruct69.52
7Llama 3.3 70B Instruct69.47
8Llama-3.1-Tulu-3-8B-DPO68.74
9Llama-3.1-Tulu-3-8B68.67
10Mistral 7B Instruct (v0.3)66.55
11Llama 3.1 70B Instruct65.3
12Llama-3.1-Tulu-3-8B-SFT64.71
13Llama 3.1 405B Instruct64.18
14Llama 3.1 8B Instruct61.3
15Llama 3.1 70B60.21

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-intent-kuq-ambiguous-f1 · How It Works · Data refreshed daily, snapshot 2026-09-19.