TW-LegalBench - Essay Questions: leaderboard

Metric: Rubric points judged fully correct (out of 580) across the 117 bar and judicial examination essay questions, zero-shot chain-of-thought; each official rubric point is labelled correct, partial, wrong or missed by gpt-5 and claude-sonnet-4.5 judges. Source: arxiv.org. Saturation forecast: Around April 2027. 13 models tracked.

Top models

#ModelScore
1GPT-5166
2Claude Sonnet 4.5160
3GPT-5.2160
4Qwen 3 235B A22B115
5GPT-OSS-120B53
6GPT-4o (2024-08-06)52
7nemotron-3-nano-30B-a3B39
8Llama 3.1 405B21
9GPT-OSS-20B17
10Qwen 2.5 7B17

Interactive version: theaggregate.ai/benchmark?slug=tw-legalbench-essay-questions · How It Works · Data refreshed daily, snapshot 2026-09-29.