AbstentionBench - underspecified context - GSM8K - Recall: leaderboard

Metric: Abstention Recall (%). Source: github.com. 23 models tracked.

Top models

#ModelScore
1O1 (Low)95.63
2O1 (High)95.47
3O195.38
4Llama 3.1 405B Instruct95.14
5Qwen 2.5 32B Instruct94.48
6Llama 3.1 70B Instruct94.31
7Gemini 1.5 Pro94.23
8Llama 3.1 8B Instruct92.66
9GPT-4o92
10Llama 3.3 70B Instruct91.1
11Llama-3.1-Tulu-3-70B-DPO89.12
12Llama-3.1-Tulu-3-70B87.8
13Mistral 7B Instruct (v0.3)87.06
14Llama-3.1-Tulu-3-8B-DPO75.35
15Llama-3.1-Tulu-3-8B74.53

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-context-gsm8k-recall · How It Works · Data refreshed daily, snapshot 2026-09-19.