HELM v2 Lite - LegalBench — leaderboard

Metric: Exact Match (%). Source: nlp.stanford.edu. 21 models tracked.

Top models

#ModelScore
1GPT-4 (0613)71.01
2Palmyra X V3 (72B)70.66
3PaLM-2 (Unicorn)70.25
4PaLM-2 (Bison)62.96
5Palmyra X V2 (33B)61.37
6Claude 261.17
7Llama 2 70B60.58
8Claude 2.160.11
9GPT-4 Preview (1106)59.26
10GPT-3.5 Turbo (0613)58.24
11text-davinci-00356.06
12text-davinci-00251.44
13Yi 34B (Base)50.66
14Llama 2 13B48.51
15LLaMA-65B46.01

Interactive version: theaggregate.ai/benchmark?slug=helm-v2-lite-legalbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.