AbstentionBench - underspecified context - MMLU Math - F1: leaderboard

Metric: Abstention F1 (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1Llama 3.1 70B Instruct84.48
2O1 (Low)83.33
3GPT-4o82.25
4O181.94
5Llama 3.1 405B Instruct81.78
6O1 (High)81.25
7Qwen 2.5 32B Instruct78.18
8Llama 3.1 8B Instruct78.03
9Llama-3.1-Tulu-3-70B-DPO76.15
10Mistral 7B Instruct (v0.3)71
11Gemini 1.5 Pro67.66
12Llama 3.3 70B Instruct64.97
13DeepSeek R1 Distill Llama 70B49.45
14Llama-3.1-Tulu-3-8B-DPO46.33
15Llama-3.1-Tulu-3-8B-SFT43.27

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-context-mmlu-math-f1 · How It Works · Data refreshed daily, snapshot 2026-09-19.