SEAL - Math — leaderboard

Metric: Score. Source: scale.com. 16 models tracked.

Top models

#ModelScore
1Claude 3.5 Sonnet (20240620)96.6
2GPT-4o (2024-08-06)95.68
3Llama 3.1 405B Instruct95.6
4Claude 3 Opus95.19
5GPT-4 Turbo (Preview)95.1
6GPT-4o (2024-05-13)94.85
7Mistral Large 2 (Jul)93.94
8Claude 3 Sonnet93.28
9Gemini 1.5 Pro (0514)92.28
10Gemini 1.5 Flash90.12
11Llama 3 70B Instruct90.12
12Mistral Large87.47
13Gemini 1.0 Pro79.83
14CodeLlama-34B-Instruct37.51

Interactive version: theaggregate.ai/benchmark?slug=seal-math · How the rankings work · Data refreshed daily, snapshot 2026-07-22.