LegalCiteBench - Case Verification: leaderboard

Metric: Cat4-2 judge score (0-100) for verifying whether a cited case supports a stated legal principle and supplying the correct case when it does not (6,377 instances), closed-book (no retrieval) on instances derived from 1,000 U.S. judicial opinions of the Case Law Access Project, greedy decoding at temperature 0 with at most 400 generated tokens, GPT-4o-mini judge; higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 21 models tracked.

Top models

#ModelScore
1Mistral 7B96.07
2DeepSeek-R1-Distill-Qwen-7B91.44
3DeepSeek R1 Distill Qwen 1.5B90.81
4Gemini 2.5 Flash89.88
5Qwen 3 14B89.58
6Qwen 2.5 7B89.32
7Llama 3.1 70B89
8GPT-5 Mini88.28
9Gemma 3 12B85.91
10O4 Mini80.12
11DeepSeek V3.178.16
12Llama 3.2 3B78.1
13Qwen 3 30B A3B78.04
14Llama 3.1 8B76.31
15Phi-474.82

Interactive version: theaggregate.ai/benchmark?slug=legalcitebench-case-verification · How It Works · Data refreshed daily, snapshot 2026-10-07.