AbstentionBench - underspecified context - MMLU History - F1: leaderboard

Metric: Abstention F1 (%). Source: github.com. 21 models tracked.

Top models

#ModelScore
1Gemini 1.5 Pro71.88
2GPT-4o67.8
3Qwen 2.5 32B Instruct62.3
4O160.71
5Llama-3.1-Tulu-3-70B-DPO58.18
6Llama 3.1 8B Instruct53.57
7Llama-3.1-Tulu-3-70B52.83
8O1 (High)52.83
9O1 (Low)52.83
10Llama 3.1 70B Instruct51.85
11Llama 3.1 405B Instruct47.27
12TinyLlama-1.1B-Chat-v1.034.62
13Llama 3.3 70B Instruct33.33
14Llama-3.1-Tulu-3-8B26.67
15Llama-3.1-Tulu-3-8B-DPO26.67

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-context-mmlu-history-f1 · How It Works · Data refreshed daily, snapshot 2026-09-19.