AbstentionBench - underspecified context - Musique - F1: leaderboard

Metric: Abstention F1 (%). Source: github.com. 20 models tracked.

Top models

#ModelScore
1Llama 3.3 70B Instruct83.08
2Llama 3.1 405B Instruct82.73
3Llama 3.1 8B Instruct82.38
4Gemini 1.5 Pro82.23
5Qwen 2.5 32B Instruct81.51
6Llama 3.1 70B Instruct80.9
7Mistral 7B Instruct (v0.3)77.45
8GPT-4o74.52
9Llama 3.1 8B70.55
10Llama 3.1 70B67.87
11Llama-3.1-Tulu-3-70B-DPO67.42
12O163.39
13Llama-3.1-Tulu-3-70B61.67
14DeepSeek R1 Distill Llama 70B51.46
15Llama-3.1-Tulu-3-8B-DPO31.04

Interactive version: theaggregate.ai/benchmark?slug=abstentionbench-underspecified-context-musique-f1 · How It Works · Data refreshed daily, snapshot 2026-09-19.