AbstentionBench - underspecified context - ALCUNA - F1: leaderboard

Metric: Abstention F1 (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1Qwen 2.5 32B Instruct83.25
2Llama-3.1-Tulu-3-70B-DPO81.87
3Llama-3.1-Tulu-3-70B81.75
4Gemini 1.5 Pro81.39
5Llama 3.1 8B Instruct81.04
6Llama 3.1 405B Instruct80.26
7Llama 3.3 70B Instruct79.91
8GPT-4o79.64
9Llama 3.1 70B Instruct79.41
10Mistral 7B Instruct (v0.3)78.41
11DeepSeek R1 Distill Llama 70B74.63
12Llama 3.1 70B72.47
13O165.62
14Llama-3.1-Tulu-3-8B-DPO64.4
15Llama-3.1-Tulu-3-8B62.86

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-context-alcuna-f1 · How It Works · Data refreshed daily, snapshot 2026-09-19.