LEXam: leaderboard

7,537 questions from 340 University of Zurich law exams in English and German: 2,841 open-ended questions graded by an LLM judge plus 4,696 multiple-choice items with 4 to 32 options; 2025, ICLR 2026.

Metric: Multiple-Choice Accuracy (%). Source: lexam-benchmark.github.io. Status: saturation imminent. 31 models tracked.

Top models

#ModelScore
1GPT-562.65
2Claude Sonnet 4.558.01
3Claude 3.7 Sonnet57.23
4Gemini 2.5 Pro55.72
5GPT-5 Mini54.82
6GPT-4.154.4
7GPT-4o53.13
8DeepSeek V3.2 Exp53.07
9DeepSeek R152.41
10Llama 4 Maverick49.1
11GPT-4.1 Mini48.49
12Qwen 3 235B A22B48.19
13QwQ-32B47.83
14GPT-OSS-120B47.71
15GPT-5 Nano47.11

Interactive version: theaggregate.ai/benchmark?slug=lexam · How It Works · Data refreshed daily, snapshot 2026-09-05.