AbstentionBench - underspecified context - MediQ - Recall: leaderboard

Metric: Abstention Recall (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1Mistral 7B Instruct (v0.3)41.45
2Qwen 2.5 32B Instruct27.96
3Gemini 1.5 Pro25.34
4GPT-4o23.21
5Llama 3.1 405B Instruct21.86
6Llama 3.3 70B Instruct19.3
7Llama 3.1 70B Instruct18.03
8Llama 3.1 8B Instruct13.34
9Llama 3.1 8B12.99
10O19.51
11Llama-3.1-Tulu-3-70B-DPO9.44
12Llama 3.1 70B7.74
13Llama-3.1-Tulu-3-70B7.59
14DeepSeek R1 Distill Llama 70B3.62
15Llama-3.1-Tulu-3-8B2.91

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-context-mediq-recall · How It Works · Data refreshed daily, snapshot 2026-09-19.