AbstentionBench - underspecified context - BBQ - F1: leaderboard

Metric: Abstention F1 (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1Qwen 2.5 32B Instruct95.78
2Llama 3.1 405B Instruct93.04
3Gemini 1.5 Pro92.15
4GPT-4o92.01
5Llama-3.1-Tulu-3-70B-DPO90.84
6Llama-3.1-Tulu-3-70B90.84
7Llama 3.1 70B Instruct89.96
8Mistral 7B Instruct (v0.3)88.46
9Llama 3.3 70B Instruct87.94
10Llama 3.1 8B Instruct85.05
11DeepSeek R1 Distill Llama 70B76.91
12Llama-3.1-Tulu-3-8B-DPO76.56
13Llama-3.1-Tulu-3-8B74.54
14O170.17
15Llama 3.1 70B61.67

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-context-bbq-f1 · How It Works · Data refreshed daily, snapshot 2026-09-19.