DLawBench: leaderboard

Metric: Resolution score (0-100), a rate scaled by 100: mean of fact resolution (annotated facts carried into the legal memo and correctly preserved or reframed against the court record) and issue resolution (expert-specified legal analysis points addressed), averaged over the two jurisdictions, on 461 real Chinese-law and U.S.-law cases replayed as multi-turn consultations with a Claude Sonnet 4.6 client simulator in four narrative styles; judged by a GPT-5.1, Claude Opus 4.6 and Gemini 3.1 Pro panel (median; same-vendor judges recuse), empty memos count as zero; provider decoding defaults; higher is better. Source: arxiv.org. Saturation forecast: Around July 2027. 26 models tracked.

Top models

#ModelScore
1GPT-5.556.2
2GPT-5.454.6
3GPT-5.249.9
4Gemini 3.1 Pro (Preview)47.2
5Claude Opus 4.744
6Qwen 3.6 Max Preview43.5
7Claude Opus 4.643.3
8Kimi K2.642.4
9GLM-5.141.8
10DeepSeek V4 Pro41.5
11Kimi K2.539.7
12Claude Sonnet 4.637.6
13Qwen 3.6 Plus37.2
14GLM-536.5
15Grok 4.1 Fast34.9

Interactive version: theaggregate.ai/benchmark?slug=dlawbench · How It Works · Data refreshed daily, snapshot 2026-09-29.