AbstentionBench - underspecified context - WorldSense - Precision: leaderboard

Metric: Abstention Precision (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1O189.25
2Llama 3.1 405B Instruct87.48
3DeepSeek R1 Distill Llama 70B86.18
4Gemini 1.5 Pro85.81
5GPT-4o84.83
6Qwen 2.5 32B Instruct77.27
7Llama 3.1 70B76.86
8Llama 3.1 8B67.76
9Llama 3.1 70B Instruct66.7
10Llama 3.3 70B Instruct65.2
11Llama-3.1-Tulu-3-70B51.75
12Llama-3.1-Tulu-3-70B-DPO51.17
13Mistral 7B Instruct (v0.3)40.91
14Llama 3.1 8B Instruct40.34
15Llama-3.1-Tulu-3-8B39.69

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-context-worldsense-precision · How It Works · Data refreshed daily, snapshot 2026-09-19.