HELM Classic - MATH Chain-of-Thought — leaderboard

Metric: Equivalent (%). Source: crfm.stanford.edu. 69 models tracked.

Top models

#ModelScore
1GPT-3.5 Turbo (0613)71.86
2GPT-3.5 Turbo (0301)68.94
3text-davinci-00344.92
4code-davinci-00243.33
5text-davinci-00238.06
6Llama 2 70B33
7Mistral-7B-v0.129.34
8LLaMA-65B25.25
9LLaMA-30B22.29
10falcon-40B13.01
11mpt-30B12.41
12Llama 2 13B12.18
13LLaMA-13B10.98
14Llama 2 7B9.91
15gpt-neox-20B7.06

Interactive version: theaggregate.ai/benchmark?slug=helm-classic-math-chain-of-thought · How the rankings work · Data refreshed daily, snapshot 2026-07-22.