LegalCiteBench - Case Matching: leaderboard

Metric: Cat4-1 judge score (0-100) for matching an anonymized legal scenario to its source case name and citation (1,997 instances), closed-book (no retrieval) on instances derived from 1,000 U.S. judicial opinions of the Case Law Access Project, greedy decoding at temperature 0 with at most 400 generated tokens, GPT-4o-mini judge; higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 21 models tracked.

Top models

#ModelScore
1GPT-5 Mini74
2DeepSeek-R1-Distill-Qwen-7B55.53
3DeepSeek R1 Distill Qwen 1.5B53.42
4Qwen 3 14B43.26
5Qwen 2.5 7B43.24
6O4 Mini42.39
7DeepSeek V3.142.29
8Claude Sonnet 4.541.64
9Gemini 2.5 Flash41.12
10Claude Haiku 4.540.56
11Phi-440.32
12Llama 3.1 8B40.24
13Qwen 3 4B40.05
14GPT-4o Mini39.54
15Qwen 3 30B A3B39.32

Interactive version: theaggregate.ai/benchmark?slug=legalcitebench-case-matching · How It Works · Data refreshed daily, snapshot 2026-10-07.