AbstentionBench - underspecified context - MediQ - Precision: leaderboard

Metric: Abstention Precision (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1Llama-3.1-Tulu-3-8B100
2GPT-4o99.39
3Llama 3.1 405B Instruct99.35
4Llama 3.1 70B Instruct99.22
5Llama 3.1 70B99.09
6Llama-3.1-Tulu-3-70B99.07
7Llama 3.1 8B98.92
8Gemini 1.5 Pro98.62
9O198.53
10Llama 3.3 70B Instruct98.19
11Llama-3.1-Tulu-3-70B-DPO97.79
12Llama 3.1 8B Instruct97.41
13DeepSeek R1 Distill Llama 70B96.23
14Mistral 7B Instruct (v0.3)95.58
15Llama-3.1-Tulu-3-8B-SFT95.45

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-context-mediq-precision · How It Works · Data refreshed daily, snapshot 2026-09-19.