AbstentionBench - underspecified context - ALCUNA - Precision: leaderboard

Metric: Abstention Precision (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1Llama 3.1 405B Instruct93.27
2Mistral 7B Instruct (v0.3)92.89
3Llama 3.3 70B Instruct92.22
4DeepSeek R1 Distill Llama 70B91.67
5Llama 3.1 70B Instruct91.3
6Gemini 1.5 Pro89.54
7Llama-3.1-Tulu-3-8B-SFT89.29
8Llama-3.1-Tulu-3-70B-DPO87.36
9Llama-3.1-Tulu-3-70B87.18
10Llama-3.1-Tulu-3-8B87.03
11O185.97
12GPT-4o85.5
13Llama 3.1 8B Instruct84.06
14Llama-3.1-Tulu-3-8B-DPO83.46
15Qwen 2.5 32B Instruct82.45

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-context-alcuna-precision · How It Works · Data refreshed daily, snapshot 2026-09-19.