AbstentionBench - underspecified intent - SituatedQA/Geo - Precision: leaderboard

Metric: Abstention Precision (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1DeepSeek R1 Distill Llama 70B86.21
2GPT-4o69.92
3Llama 3.3 70B Instruct69.03
4Llama 3.1 70B59.57
5O159.04
6Llama 3.1 8B Instruct58.67
7Mistral 7B Instruct (v0.3)58.52
8Llama 3.1 8B55.79
9Llama 3.1 70B Instruct55.74
10Llama-3.1-Tulu-3-8B52.94
11Gemini 1.5 Pro50.35
12Llama-3.1-Tulu-3-70B-DPO50.22
13Llama-3.1-Tulu-3-70B49.78
14Llama 3.1 405B Instruct49.68
15Llama-3.1-Tulu-3-8B-DPO49.04

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-intent-situatedqa-geo-precision · How It Works · Data refreshed daily, snapshot 2026-09-19.