HELM Classic - LSAT — leaderboard

Metric: Exact Match (%). Source: crfm.stanford.edu. 69 models tracked.

Top models

#ModelScore
1GPT-3.5 Turbo (0301)25.22
2mpt-30B24.78
3Llama 2 13B23.48
4Llama 2 70B23.48
5text-davinci-00323.33
6text-davinci-00222.9
7alpaca-7B22.17
8Mistral-7B-v0.121.3
9LLaMA-65B21.3
10LLaMA-30B21.3
11falcon-40B21.3
12text-ada-00121.3
13LLaMA-7B20.87
14falcon-7B20.43
15pythia-6.9B19.57

Interactive version: theaggregate.ai/benchmark?slug=helm-classic-lsat · How the rankings work · Data refreshed daily, snapshot 2026-07-22.