LEXam — leaderboard

LEXam evaluates model capability on legal tasks from the linked upstream source with Open Question Judge Score as the primary reported metric.

Metric: Multiple-Choice Accuracy (%). Source: lexam-benchmark.github.io. Status: saturation imminent. 31 models tracked.

Top models

#ModelScore
1GPT-562.65
2Claude Sonnet 4.558.01
3Claude 3.7 Sonnet57.23
4Gemini 2.5 Pro55.72
5GPT-5 Mini54.82
6GPT-4.154.4
7GPT-4o53.13
8DeepSeek V3.2 Exp53.07
9DeepSeek R152.41
10Llama 4 Maverick49.1
11GPT-4.1 Mini48.49
12Qwen 3 235B A22B48.19
13QwQ-32B47.83
14GPT-OSS-120B47.71
15GPT-5 Nano47.11

Interactive version: theaggregate.ai/benchmark?slug=lexam · How the rankings work · Data refreshed daily, snapshot 2026-07-22.