AbstentionBench - underspecified context - GSM8K - Precision: leaderboard

Metric: Abstention Precision (%). Source: github.com. 23 models tracked.

Top models

#ModelScore
1Llama 3.3 70B Instruct100
2GPT-4o100
3Qwen 2.5 32B Instruct100
4Llama-3.1-Tulu-3-8B100
5Llama-3.1-Tulu-3-70B100
6Llama-3.1-Tulu-3-70B-DPO100
7O1 (High)100
8O1 (Low)99.91
9O199.91
10Llama 3.1 70B Instruct99.91
11Gemini 1.5 Pro99.91
12Llama-3.1-Tulu-3-8B-DPO99.78
13Llama 3.1 8B Instruct99.73
14Llama 3.1 405B Instruct99.65
15Llama-3.1-Tulu-3-8B-SFT99.52

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-context-gsm8k-precision · How It Works · Data refreshed daily, snapshot 2026-09-19.