TW-LegalBench - Judgment Prediction: leaderboard

Metric: ROUGE-L x100 (0-100), Jieba-segmented, of the predicted judgment holding against the court's main result, 535 first-instance criminal judgments over 107 offense categories, zero-shot chain-of-thought. Source: arxiv.org. Saturation forecast: Around April 2028. 12 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.553.5
2GPT-5.242.7
3GPT-539.9
4Qwen 3 235B A22B32.5
5Llama 3.1 405B29.8
6GPT-4o (2024-08-06)28.4
7GPT-OSS-120B26.3
8GPT-OSS-20B24.6
9nemotron-3-nano-30B-a3B22.5
10Qwen 2.5 7B21.7

Interactive version: theaggregate.ai/benchmark?slug=tw-legalbench-judgment-prediction · How It Works · Data refreshed daily, snapshot 2026-09-29.