AbstentionBench - false premise - FalseQA - F1: leaderboard

Metric: Abstention F1 (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1Gemini 1.5 Pro75.54
2Qwen 2.5 32B Instruct74.51
3O174.26
4GPT-4o72.38
5Llama 3.3 70B Instruct71.96
6Llama 3.1 405B Instruct68.37
7Llama 3.1 70B Instruct66.79
8Llama-3.1-Tulu-3-70B-DPO64.96
9Mistral 7B Instruct (v0.3)64.89
10Llama-3.1-Tulu-3-70B64.83
11Llama 3.1 8B Instruct60.74
12DeepSeek R1 Distill Llama 70B58.74
13Llama-3.1-Tulu-3-8B-SFT57.17
14Llama-3.1-Tulu-3-8B-DPO55.8
15Llama-3.1-Tulu-3-8B54.8

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-false-premise-falseqa-f1 · How It Works · Data refreshed daily, snapshot 2026-09-19.