AbstentionBench - underspecified context - Musique - Precision: leaderboard

Metric: Abstention Precision (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1Llama 3.1 70B99.41
2Llama 3.1 8B99.31
3Llama-3.1-Tulu-3-8B98.73
4Llama-3.1-Tulu-3-8B-SFT97.06
5Llama-3.1-Tulu-3-8B-DPO95.31
6GPT-4o91.68
7O191.29
8Llama 3.1 70B Instruct90.11
9Llama 3.1 405B Instruct89.41
10Llama-3.1-Tulu-3-70B-DPO89.24
11Llama-3.1-Tulu-3-70B88.97
12DeepSeek R1 Distill Llama 70B88.66
13Llama 3.3 70B Instruct88.42
14Llama 3.1 8B Instruct83.98
15Gemini 1.5 Pro80.99

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-context-musique-precision · How It Works · Data refreshed daily, snapshot 2026-09-19.