LitBench - Citation Recommendation (Quantitative Biology): leaderboard

Metric: Accuracy (%) at picking the paper actually cited, out of 11 candidates (the cited paper and 10 random papers of the same domain graph), for citation edges of LitBench's held-out quantitative-biology citation subgraph built from arXiv; greedy decoding; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 7 models tracked.

Top models

#ModelScoreOverall rank
1DeepSeek R173.3#245
2GPT-4o71.32#333
3Mistral 7B27.94#1377
4Llama 3 8B19.23#1231
5Llama 3.2 3B13.6#1384
6Vicuna-7B12.5#1521
7Llama 3.2 1B4.4#1549

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=litbench-citation-recommendation-quantitative-biology · How It Works · Data refreshed daily, snapshot 2026-10-11.