LegalCiteBench - Misleading Answer Rate: leaderboard

Metric: Misleading Answer Rate (%): share of low-scoring responses (task score under 40) on Cat1, Cat2 and Cat4-1 that still give a concrete citation or case instead of abstaining, closed-book (no retrieval) on instances derived from 1,000 U.S. judicial opinions of the Case Law Access Project, greedy decoding at temperature 0 with at most 400 generated tokens, GPT-4o-mini judge; lower is better. Source: arxiv.org. Saturation forecast: Around 2040. 21 models tracked.

Top models

#ModelScore
1GPT-5 Mini89.24
2DeepSeek-R1-Distill-Qwen-7B94.21
3DeepSeek R1 Distill Qwen 1.5B94.62
4Claude Sonnet 4.596.79
5O4 Mini96.85
6Qwen 3 14B96.97
7Qwen 2.5 7B97.53
8DeepSeek V3.197.58
9Gemini 2.5 Flash98.21
10Claude Haiku 4.598.31
11Qwen 3 30B A3B98.35
12Llama 3.1 70B98.47
13Mistral 7B98.69
14Qwen 3 4B99.01
15Phi-499.1

Interactive version: theaggregate.ai/benchmark?slug=legalcitebench-misleading-answer-rate · How It Works · Data refreshed daily, snapshot 2026-10-07.