AbstentionBench - answer unknown - CoCoNot/Unsupported - Precision: leaderboard

Metric: Abstention Precision (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1Mistral 7B Instruct (v0.3)100
2O1100
3DeepSeek R1 Distill Llama 70B100
4Llama-3.1-Tulu-3-70B-DPO99.14
5Qwen 2.5 32B Instruct99.11
6Llama-3.1-Tulu-3-8B-DPO99.11
7GPT-4o99.1
8Llama-3.1-Tulu-3-8B98.21
9Llama 3.1 405B Instruct98.13
10Llama 3.1 8B Instruct98.11
11Llama 3.1 70B Instruct98.06
12Gemini 1.5 Pro97.41
13Llama-3.1-Tulu-3-70B96.64
14Llama 3.3 70B Instruct96.26
15Llama-3.1-Tulu-3-8B-SFT95.08

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-answer-unknown-coconot-unsupported-precision · How It Works · Data refreshed daily, snapshot 2026-09-19.