TW-LegalBench - Administrative Law: leaderboard

Metric: Administrative-law questions: accuracy (%) on four-option single-answer questions from Taiwan's 2020-2024 official legal examinations (civil service, judicial, police and professional), zero-shot chain-of-thought, JSON answer (format failures count as wrong), temperature 0 (1 for gpt-5 and gpt-5.2). Source: arxiv.org. Saturation forecast: Around December 2026. 13 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.578.1
2GPT-572.3
3GPT-5.269.7
4Qwen 3 235B A22B64.4
5GPT-4o (2024-08-06)61.5
6Llama 3.1 405B56.6
7GPT-OSS-120B52.1
8nemotron-3-nano-30B-a3B48.7
9Qwen 2.5 7B46.7
10GPT-OSS-20B44.9

Interactive version: theaggregate.ai/benchmark?slug=tw-legalbench-administrative-law · How It Works · Data refreshed daily, snapshot 2026-09-29.