LegalCiteTrust: leaderboard
Metric: Final score (0 to 2; per report (Coverage + Support) x Trust, each dimension in 0-1, macro-averaged over the 72 Chinese legal research tasks; Coverage of the rubric issues, Support by statute and case evidence, Trust as the mean product of Existence, Fidelity and Applicability of each cited law or case; each system in its own usage mode, direct LLMs without tools). Source: arxiv.org. Saturation forecast: Around April 2027. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5 | 0.98 |
| 2 | Qwen 3.6 Plus | 0.96 |
Interactive version: theaggregate.ai/benchmark?slug=legalcitetrust · How It Works · Data refreshed daily, snapshot 2026-09-29.