TW-LegalBench - Civil Law: leaderboard

Metric: Civil-law questions: accuracy (%) on four-option single-answer questions from Taiwan's 2020-2024 official legal examinations (civil service, judicial, police and professional), zero-shot chain-of-thought, JSON answer (format failures count as wrong), temperature 0 (1 for gpt-5 and gpt-5.2). Source: arxiv.org. Saturation forecast: Around December 2026. 13 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.580.9
2GPT-577.7
3GPT-5.274
4Qwen 3 235B A22B69.1
5GPT-4o (2024-08-06)64.9
6Llama 3.1 405B57.6
7GPT-OSS-120B52.7
8nemotron-3-nano-30B-a3B49.2
9Qwen 2.5 7B46.8
10GPT-OSS-20B44.4

Interactive version: theaggregate.ai/benchmark?slug=tw-legalbench-civil-law · How It Works · Data refreshed daily, snapshot 2026-09-29.