TW-LegalBench - International Law: leaderboard

Metric: International-law questions: accuracy (%) on four-option single-answer questions from Taiwan's 2020-2024 official legal examinations (civil service, judicial, police and professional), zero-shot chain-of-thought, JSON answer (format failures count as wrong), temperature 0 (1 for gpt-5 and gpt-5.2). Source: arxiv.org. Saturation forecast: Around December 2026. 13 models tracked.

Top models

#ModelScore
1GPT-5.276.6
2GPT-575.8
3Claude Sonnet 4.575.8
4Qwen 3 235B A22B69.4
5GPT-4o (2024-08-06)67.7
6Llama 3.1 405B67.7
7GPT-OSS-120B62.1
8nemotron-3-nano-30B-a3B56.5
9GPT-OSS-20B55.6
10Qwen 2.5 7B53.2

Interactive version: theaggregate.ai/benchmark?slug=tw-legalbench-international-law · How It Works · Data refreshed daily, snapshot 2026-09-29.