HELM Classic - Synthetic Reasoning Abstract — leaderboard

Metric: Exact Match (%). Source: crfm.stanford.edu. 69 models tracked.

Top models

#ModelScore
1code-davinci-00253.96
2GPT-3.5 Turbo (0613)50.94
3text-davinci-00350.25
4Llama 2 70B48.87
5text-davinci-00248.78
6Mistral-7B-v0.145.11
7GPT-3.5 Turbo (0301)44.98
8LLaMA-65B43.82
9LLaMA-30B41.23
10mpt-30B33.53
11Llama 2 13B32.36
12falcon-40B23.75
13davinci23.58
14Llama 2 7B22.2
15LLaMA-13B21.42

Interactive version: theaggregate.ai/benchmark?slug=helm-classic-synthetic-reasoning-abstract · How the rankings work · Data refreshed daily, snapshot 2026-07-22.