FDARxBench - Refusal F1: leaderboard

Metric: Macro F1 (x 100) for abstaining on unanswerable questions (an out-of-label entity inserted) versus answering answerable ones, from FDARxBench's expert-guided QA items grounded in 700 FDA prescription drug labels, where the whole drug label is given as passage-indexed context and the model answers with cited passage ids; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 10 models tracked.

Top models

#ModelScoreOverall rank
1Llama 3.1 8B Instruct79.6#1018
2Llama 3.3 70B Instruct78.9#520
3Qwen 3 32B75.6#424
4Claude Sonnet 4.574.8#138
5Claude Opus 4.673.1#60
6GPT-5.172.5#131
7GPT-4o Mini71.7#588
8Ministral-3-14B-Instruct-251271#590
9Qwen 3 14B70.9#524
10GPT-5.270.1#105

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=fdarxbench-refusal-f1 · How It Works · Data refreshed daily, snapshot 2026-10-11.