AbstentionBench - underspecified intent - CoCoNot/Incomprehensible - F1: leaderboard

Metric: Abstention F1 (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1Qwen 2.5 32B Instruct97.92
2Llama-3.1-Tulu-3-8B-SFT97.92
3GPT-4o96.84
4Llama-3.1-Tulu-3-70B-DPO96.84
5Llama 3.1 8B Instruct95.74
6Llama-3.1-Tulu-3-8B95.74
7Gemini 1.5 Pro94.62
8Llama-3.1-Tulu-3-70B94.62
9Llama-3.1-Tulu-3-8B-DPO93.48
10Llama 3.1 405B Instruct92.31
11O192.31
12Llama 3.1 70B Instruct91.11
13Mistral 7B Instruct (v0.3)91.11
14Llama 3.3 70B Instruct83.33
15DeepSeek R1 Distill Llama 70B75.95

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-intent-coconot-incomprehensible-f1 · How It Works · Data refreshed daily, snapshot 2026-09-19.