AbstentionBench - underspecified context - MMLU Math - Precision: leaderboard

Metric: Abstention Precision (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1Llama 3.3 70B Instruct100
2Llama 3.1 405B Instruct100
3Gemini 1.5 Pro100
4O1 (High)100
5O1 (Low)100
6Llama 3.1 70B Instruct98.99
7O198.94
8Qwen 2.5 32B Instruct98.85
9Llama-3.1-Tulu-3-70B-DPO97.65
10Llama-3.1-Tulu-3-8B-SFT97.37
11GPT-4o96.94
12Llama-3.1-Tulu-3-8B96.77
13Llama 3.1 8B Instruct96.67
14Llama-3.1-Tulu-3-8B-DPO93.18
15DeepSeek R1 Distill Llama 70B91.84

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-context-mmlu-math-precision · How It Works · Data refreshed daily, snapshot 2026-09-19.