HELM Classic - MATH — leaderboard

Metric: Equivalent (%). Source: crfm.stanford.edu. 69 models tracked.

Top models

#ModelScore
1GPT-3.5 Turbo (0301)48.83
2GPT-3.5 Turbo (0613)45.28
3code-davinci-00241.02
4text-davinci-00339.01
5text-davinci-00232.79
6Llama 2 70B26.08
7LLaMA-65B22.4
8falcon-40B20.98
9Mistral-7B-v0.120.87
10LLaMA-30B19.66
11mpt-30B17.81
12Llama 2 13B14.46
13gpt-neox-20B14.05
14LLaMA-13B13.36
15LLaMA-7B11.19

Interactive version: theaggregate.ai/benchmark?slug=helm-classic-math · How the rankings work · Data refreshed daily, snapshot 2026-07-22.