DLawBench - Elicitation: leaderboard

Metric: Elicitation score (0-100), a rate scaled by 100: mean of fact coverage (annotated case facts the lawyer collects) and inquiry (expert-specified follow-up questions asked), averaged over the two jurisdictions, on 461 real Chinese-law and U.S.-law cases replayed as multi-turn consultations with a Claude Sonnet 4.6 client simulator in four narrative styles; judged by a GPT-5.1, Claude Opus 4.6 and Gemini 3.1 Pro panel (median; same-vendor judges recuse), empty memos count as zero; provider decoding defaults; higher is better. Source: arxiv.org. Saturation forecast: Around October 2027. 26 models tracked.

Top models

#ModelScore
1GPT-5.570.7
2GPT-5.468
3GPT-5.267.5
4Grok 4.1 Fast59.6
5Seed 2.0 Pro57
6Claude Opus 4.656.6
7Gemini 3.1 Pro (Preview)56.5
8GLM-5.156.3
9Qwen 3.6 Max Preview55.5
10Kimi K2.655.3
11DeepSeek V4 Pro54.5
12Claude Opus 4.754.3
13Kimi K2.553.5
14GLM-552.1
15Qwen 3.6 Plus51.5

Interactive version: theaggregate.ai/benchmark?slug=dlawbench-elicitation · How It Works · Data refreshed daily, snapshot 2026-09-29.