AbstentionBench - false premise - CoCoNot/False presumptions - F1: leaderboard

Metric: Abstention F1 (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1O189.57
2Llama-3.1-Tulu-3-70B88.16
3Llama-3.1-Tulu-3-70B-DPO86.84
4GPT-4o84.35
5Llama 3.1 405B Instruct83.54
6Qwen 2.5 32B Instruct83.22
7Llama 3.3 70B Instruct82.99
8Llama-3.1-Tulu-3-8B-DPO82.89
9Llama-3.1-Tulu-3-8B81.63
10Mistral 7B Instruct (v0.3)80.82
11Llama 3.1 70B Instruct80
12DeepSeek R1 Distill Llama 70B79.14
13Gemini 1.5 Pro78.91
14Llama-3.1-Tulu-3-8B-SFT77.37
15Llama 3.1 8B Instruct76.51

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-false-premise-coconot-false-presumptions-f1 · How It Works · Data refreshed daily, snapshot 2026-09-19.