LegalCiteTrust: leaderboard

Metric: Final score (0 to 2; per report (Coverage + Support) x Trust, each dimension in 0-1, macro-averaged over the 72 Chinese legal research tasks; Coverage of the rubric issues, Support by statute and case evidence, Trust as the mean product of Existence, Fidelity and Applicability of each cited law or case; each system in its own usage mode, direct LLMs without tools). Source: arxiv.org. Saturation forecast: Around April 2027. 7 models tracked.

Top models

#ModelScore
1GPT-50.98
2Qwen 3.6 Plus0.96

Interactive version: theaggregate.ai/benchmark?slug=legalcitetrust · How It Works · Data refreshed daily, snapshot 2026-09-29.