HELM Lite - MATH Level 1 - Precalculus: leaderboard

Metric: Equivalent (CoT) (%). Source: crfm.stanford.edu. 91 models tracked.

Top models

#ModelScore
1Gemini 2.0 Flash (Preview)96.49
2Gemini 1.5 Pro (002)94.74
3Gemini 1.5 Flash (002)91.23
4DeepSeek V391.23
5Claude 3.5 Sonnet (20241022)87.72
6Nova Pro84.21
7Qwen 2.5 72B Instruct82.46
8GPT-4o (2024-08-06)82.46
9Gemma 2 27B (IT)82.46
10Qwen 2 72B Instruct80.7
11Llama 3.1 405B Instruct78.95
12Claude 3.5 Haiku (20241022)78.95
13Gemini 1.5 Pro (001)78.95
14Gemini 1.5 Flash (001)78.95
15Nova Lite78.95

Interactive version: theaggregate.ai/benchmark?slug=helm-lite-math-level-1-precalculus · How It Works · Data refreshed daily, snapshot 2026-09-19.