LegalCiteBench - Citation Error Detection: leaderboard

Metric: Cat3 judge score (0-100) for detecting citation errors in legal analysis paragraphs and supplying the correct citation (5,474 instances), closed-book (no retrieval) on instances derived from 1,000 U.S. judicial opinions of the Case Law Access Project, greedy decoding at temperature 0 with at most 400 generated tokens, GPT-4o-mini judge; higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 21 models tracked.

Top models

#ModelScore
1DeepSeek-R1-Distill-Qwen-7B58.71
2DeepSeek R1 Distill Qwen 1.5B58.11
3GPT-5 Mini54.95
4Qwen 3 14B52.18
5Qwen 3 4B47.12
6Gemini 2.5 Flash46.93
7O4 Mini46.71
8Llama 3.1 8B45.32
9Claude Sonnet 4.544.87
10Qwen 3 30B A3B43.55
11Mistral 7B43.4
12Phi-443.35
13Claude Haiku 4.543.01
14GPT-4o Mini42.84
15Llama 3.2 3B42.27

Interactive version: theaggregate.ai/benchmark?slug=legalcitebench-citation-error-detection · How It Works · Data refreshed daily, snapshot 2026-10-07.