FDARxBench - Citation F1 (Multi-Hop): leaderboard

Metric: Macro F1 (x 100) between cited and gold provenance passage ids on multi-hop questions from FDARxBench's expert-guided QA items grounded in 700 FDA prescription drug labels, where the whole drug label is given as passage-indexed context and the model answers with cited passage ids; higher is better. Source: arxiv.org. Saturation forecast: Around 2034. 10 models tracked.

Top models

#ModelScoreOverall rank
1Claude Opus 4.645.8#60
2Llama 3.3 70B Instruct43.3#520
3Qwen 3 32B43.3#424
4GPT-5.140.6#131
5Claude Sonnet 4.538.3#138
6GPT-5.238.3#105
7GPT-4o Mini37.8#588
8Qwen 3 14B36.2#524
9Llama 3.1 8B Instruct33.3#1018
10Ministral-3-14B-Instruct-251232.8#590

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=fdarxbench-citation-f1-multi-hop · How It Works · Data refreshed daily, snapshot 2026-10-11.