LegalCiteBench - Citation Completion: leaderboard

Metric: Cat2 citation-level F1 (0-100) of the remaining source-opinion authorities the model supplies given a partial citation set (4,899 instances), closed-book (no retrieval) on instances derived from 1,000 U.S. judicial opinions of the Case Law Access Project, greedy decoding at temperature 0 with at most 400 generated tokens, GPT-4o-mini judge; higher is better. Source: arxiv.org. Saturation forecast: Around 2039. 21 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.56.35
2Qwen 3 14B6.34
3DeepSeek V3.15.15
4DeepSeek R1 Distill Qwen 1.5B4.54
5Qwen 3 30B A3B4.34
6Llama 3.1 70B4.34
7O4 Mini4.06
8Claude Haiku 4.53.87
9Gemini 2.5 Flash3.66
10Qwen 2.5 7B3.5
11DeepSeek-R1-Distill-Qwen-7B3.49
12Qwen 3 4B3.37
13Phi-43.16
14Mistral 7B2.56
15GPT-4o Mini2.27

Interactive version: theaggregate.ai/benchmark?slug=legalcitebench-citation-completion · How It Works · Data refreshed daily, snapshot 2026-10-07.