AbstentionBench - underspecified intent - CoCoNot/Incomprehensible - Recall: leaderboard

Metric: Abstention Recall (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1Qwen 2.5 32B Instruct95.92
2Llama-3.1-Tulu-3-8B-SFT95.92
3GPT-4o93.88
4Llama-3.1-Tulu-3-70B-DPO93.88
5Llama 3.1 8B Instruct91.84
6Llama-3.1-Tulu-3-8B91.84
7Gemini 1.5 Pro89.8
8Llama-3.1-Tulu-3-70B89.8
9Llama-3.1-Tulu-3-8B-DPO87.76
10Llama 3.1 405B Instruct85.71
11O185.71
12Llama 3.1 70B Instruct83.67
13Mistral 7B Instruct (v0.3)83.67
14Llama 3.3 70B Instruct71.43
15DeepSeek R1 Distill Llama 70B61.22

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-intent-coconot-incomprehensible-recall · How It Works · Data refreshed daily, snapshot 2026-09-19.