DLawBench - Chinese Law: leaderboard

Metric: Resolution score (0-100), a rate scaled by 100: mean of fact resolution (annotated facts carried into the legal memo and correctly preserved or reframed against the court record) and issue resolution (expert-specified legal analysis points addressed), Chinese-law cases only, on 461 real Chinese-law and U.S.-law cases replayed as multi-turn consultations with a Claude Sonnet 4.6 client simulator in four narrative styles; judged by a GPT-5.1, Claude Opus 4.6 and Gemini 3.1 Pro panel (median; same-vendor judges recuse), empty memos count as zero; provider decoding defaults; higher is better. Source: arxiv.org. Saturation forecast: Around March 2027. 26 models tracked.

Top models

#ModelScore
1GPT-5.561.2
2GPT-5.459.5
3GPT-5.255.4
4Gemini 3.1 Pro (Preview)51
5Kimi K2.650
6Qwen 3.6 Max Preview48
7Claude Opus 4.746.8
8Kimi K2.546
9GLM-5.144.3
10Claude Opus 4.643.8
11DeepSeek V4 Pro43.6
12Seed 2.0 Pro42.5
13Qwen 3.6 Plus42.2
14GLM-538.8
15DeepSeek V3.2 (Thinking)38.5

Interactive version: theaggregate.ai/benchmark?slug=dlawbench-chinese-law · How It Works · Data refreshed daily, snapshot 2026-09-29.