AbstentionBench - underspecified context - MediQ - F1: leaderboard

Metric: Abstention F1 (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1Mistral 7B Instruct (v0.3)57.82
2Qwen 2.5 32B Instruct43.25
3Gemini 1.5 Pro40.32
4GPT-4o37.63
5Llama 3.1 405B Instruct35.83
6Llama 3.3 70B Instruct32.27
7Llama 3.1 70B Instruct30.51
8Llama 3.1 8B Instruct23.47
9Llama 3.1 8B22.96
10O117.35
11Llama-3.1-Tulu-3-70B-DPO17.22
12Llama 3.1 70B14.35
13Llama-3.1-Tulu-3-70B14.11
14DeepSeek R1 Distill Llama 70B6.98
15Llama-3.1-Tulu-3-8B5.66

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-context-mediq-f1 · How It Works · Data refreshed daily, snapshot 2026-09-19.