DLawBench - US Law: leaderboard

Metric: Resolution score (0-100), a rate scaled by 100: mean of fact resolution (annotated facts carried into the legal memo and correctly preserved or reframed against the court record) and issue resolution (expert-specified legal analysis points addressed), U.S.-law cases only, on 461 real Chinese-law and U.S.-law cases replayed as multi-turn consultations with a Claude Sonnet 4.6 client simulator in four narrative styles; judged by a GPT-5.1, Claude Opus 4.6 and Gemini 3.1 Pro panel (median; same-vendor judges recuse), empty memos count as zero; provider decoding defaults; higher is better. Source: arxiv.org. Saturation forecast: Around September 2027. 26 models tracked.

Top models

#ModelScore
1GPT-5.551.2
2GPT-5.449.7
3GPT-5.244.4
4Gemini 3.1 Pro (Preview)43.3
5Claude Opus 4.642.9
6Claude Opus 4.741.2
7DeepSeek V4 Pro39.4
8GLM-5.139.3
9Qwen 3.6 Max Preview39.1
10Claude Sonnet 4.637.8
11Grok 4.1 Fast35.9
12Kimi K2.634.7
13GLM-534.2
14Kimi K2.533.4
15Qwen 3.6 Plus32.2

Interactive version: theaggregate.ai/benchmark?slug=dlawbench-us-law · How It Works · Data refreshed daily, snapshot 2026-09-29.