HELM Lite - LegalBench: leaderboard

Metric: Exact Match (%). Source: crfm.stanford.edu. 91 models tracked.

Top models

#ModelScore
1Gemini 1.5 Pro (001)75.72
2Gemini 1.5 Pro (002)74.7
3Qwen 2.5 72B Instruct74.01
4Nova Pro73.62
5GPT-4o (2024-05-13)73.32
6Llama 3 70B73.29
7GPT-4 Turbo72.73
8Llama 3.3 70B Instruct72.54
9GPT-4o (2024-08-06)72.09
10DeepSeek V371.78
11GPT-4 (0613)71.28
12Qwen 2 72B Instruct71.17
13Mixtral 8x22B70.79
14Llama 3.1 405B Instruct70.71
15Claude 3.5 Sonnet (20240620)70.66

Interactive version: theaggregate.ai/benchmark?slug=helm-lite-legalbench · How It Works · Data refreshed daily, snapshot 2026-09-19.