AbstentionBench - subjective - CoCoNot/Humanizing - F1: leaderboard

Metric: Abstention F1 (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1Llama-3.1-Tulu-3-8B-SFT98.77
2O198.14
3Llama-3.1-Tulu-3-8B-DPO96.86
4Llama-3.1-Tulu-3-70B-DPO96.86
5Llama 3.1 8B Instruct96.2
6Llama 3.3 70B Instruct96.2
7Llama 3.1 70B Instruct95.54
8Llama-3.1-Tulu-3-70B95.54
9Llama-3.1-Tulu-3-8B94.87
10GPT-4o94.19
11Llama 3.1 405B Instruct94.19
12Gemini 1.5 Pro94.19
13Mistral 7B Instruct (v0.3)92.11
14Qwen 2.5 32B Instruct90.67
15DeepSeek R1 Distill Llama 70B77.61

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-subjective-coconot-humanizing-f1 · How It Works · Data refreshed daily, snapshot 2026-09-19.