HELM Lite - MATH Level 1 - Counting and Probability: leaderboard

Metric: Equivalent (CoT) (%). Source: crfm.stanford.edu. 91 models tracked.

Top models

#ModelScore
1Gemini 1.5 Flash (002)97.44
2Gemini 1.5 Pro (002)94.87
3Claude 3.5 Sonnet (20241022)92.31
4Qwen 2.5 72B Instruct92.31
5GPT-4o (2024-08-06)92.31
6GPT-4o (2024-05-13)92.31
7GPT-4 Turbo92.31
8Gemini 2.0 Flash (Preview)92.31
9GPT-4 Turbo (Preview)92.31
10DeepSeek V389.74
11Llama 3.2 90B Vision Instruct89.74
12Claude 3.5 Haiku (20241022)89.74
13Llama 3.3 70B Instruct87.18
14Llama 3.1 70B Instruct87.18
15GPT-4o Mini (2024-07-18)87.18

Interactive version: theaggregate.ai/benchmark?slug=helm-lite-math-level-1-counting-and-probability · How It Works · Data refreshed daily, snapshot 2026-09-19.