FDARxBench - Citation F1 (Factual): leaderboard

Metric: Macro F1 (x 100) between cited and gold provenance passage ids on factual questions from FDARxBench's expert-guided QA items grounded in 700 FDA prescription drug labels, where the whole drug label is given as passage-indexed context and the model answers with cited passage ids; higher is better. Source: arxiv.org. Saturation forecast: Around 2032. 10 models tracked.

Top models

#ModelScoreOverall rank
1Claude Sonnet 4.552.8#138
2Qwen 3 32B52.5#424
3Claude Opus 4.652.2#60
4GPT-5.252#105
5GPT-5.152#131
6Llama 3.3 70B Instruct50.8#520
7Qwen 3 14B50#524
8GPT-4o Mini49.7#588
9Llama 3.1 8B Instruct45.8#1018
10Ministral-3-14B-Instruct-251243.9#590

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=fdarxbench-citation-f1-factual · How It Works · Data refreshed daily, snapshot 2026-10-11.