AbstentionBench - false premise - QAQA - F1: leaderboard

Metric: Abstention F1 (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1O174.13
2GPT-4o68.14
3Gemini 1.5 Pro67.51
4Llama-3.1-Tulu-3-70B66.39
5Llama-3.1-Tulu-3-70B-DPO63.81
6Llama 3.1 405B Instruct63
7Qwen 2.5 32B Instruct59.46
8Llama-3.1-Tulu-3-8B-DPO59.18
9Mistral 7B Instruct (v0.3)57.6
10Llama 3.3 70B Instruct55.4
11Llama-3.1-Tulu-3-8B54.75
12Llama 3.1 70B Instruct54.55
13Llama 3.1 8B Instruct53.39
14Llama-3.1-Tulu-3-8B-SFT47.8
15DeepSeek R1 Distill Llama 70B47.52

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-false-premise-qaqa-f1 · How It Works · Data refreshed daily, snapshot 2026-09-19.