DLawBench - Fidelity: leaderboard

Metric: Fidelity score (0-100), a rate scaled by 100: share of factual and inferential claims in the legal memo that the consultation supports, averaged over the two jurisdictions, on 461 real Chinese-law and U.S.-law cases replayed as multi-turn consultations with a Claude Sonnet 4.6 client simulator in four narrative styles; judged by a GPT-5.1, Claude Opus 4.6 and Gemini 3.1 Pro panel (median; same-vendor judges recuse), empty memos count as zero; provider decoding defaults; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 26 models tracked.

Top models

#ModelScore
1GPT-5.494
2GPT-5.293.6
3GPT-5.593.4
4DeepSeek V3.2 (Thinking)89.8
5Claude Sonnet 4.688.9
6GLM-5.188.4
7GLM-588
8Seed 2.0 Pro87.7
9Claude Opus 4.687.6
10Claude Opus 4.787.3
11Qwen 3.6 Max Preview87
12Kimi K2.686.8
13Qwen 3.6 Plus86.5
14Qwen 3.5 Plus85.9
15DeepSeek V4 Pro85.1

Interactive version: theaggregate.ai/benchmark?slug=dlawbench-fidelity · How It Works · Data refreshed daily, snapshot 2026-09-29.