Omni-MATH: leaderboard

Olympiad-level mathematics benchmark spanning domains and difficulty bands, using structured grading to evaluate advanced mathematical reasoning.

Metric: Overall Accuracy (%). Source: omni-math.github.io. Status: saturated. 15 models tracked.

Top models

#ModelScore
1O1 Mini60.54
2O1 Preview52.55
3Qwen2.5-Math-72B-Instruct36.2
4Qwen2.5-Math-7B-Instruct33.22
5GPT-4o30.49
6Claude 3.5 Sonnet26.23
7DeepSeek Coder V225.78
8Llama 3.1 70B Instruct24.16

Interactive version: theaggregate.ai/benchmark?slug=omni-math · How It Works · Data refreshed daily, snapshot 2026-09-05.