TW-LegalBench - Constitutional Law: leaderboard

Metric: Constitutional-law questions: accuracy (%) on four-option single-answer questions from Taiwan's 2020-2024 official legal examinations (civil service, judicial, police and professional), zero-shot chain-of-thought, JSON answer (format failures count as wrong), temperature 0 (1 for gpt-5 and gpt-5.2). Source: arxiv.org. Saturation forecast: Around December 2026. 13 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.589.4
2GPT-587.6
3GPT-5.284.8
4GPT-4o (2024-08-06)79.9
5Qwen 3 235B A22B78.5
6Llama 3.1 405B72.2
7GPT-OSS-120B67.2
8nemotron-3-nano-30B-a3B63.5
9Qwen 2.5 7B59.3
10GPT-OSS-20B59.1

Interactive version: theaggregate.ai/benchmark?slug=tw-legalbench-constitutional-law · How It Works · Data refreshed daily, snapshot 2026-09-29.