LegalCiteBench - Citation Retrieval: leaderboard

Metric: Cat1 citation-level F1 (0-100) between the cited authorities the model lists for an opinion-derived legal research question and the reporter citations of the source opinion (4,899 instances), closed-book (no retrieval) on instances derived from 1,000 U.S. judicial opinions of the Case Law Access Project, greedy decoding at temperature 0 with at most 400 generated tokens, GPT-4o-mini judge; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 21 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.56.8
2O4 Mini5.73
3DeepSeek V3.15.31
4Gemini 2.5 Flash5.2
5Claude Haiku 4.54.42
6Llama 3.1 70B3.82
7Qwen 3 30B A3B3.65
8Phi-43.59
9Qwen 3 4B3.33
10Mistral 7B2.85
11Qwen 3 14B2.54
12GPT-4o Mini2.4
13Gemma 3 12B1.92
14GPT-5 Mini1.81
15Llama 3.1 8B1.47

Interactive version: theaggregate.ai/benchmark?slug=legalcitebench-citation-retrieval · How It Works · Data refreshed daily, snapshot 2026-10-07.