IneqMath — leaderboard

Olympiad-level inequality proof benchmark evaluating both final-answer correctness and step-wise reasoning soundness for chat and reasoning LLMs.

Metric: Overall Accuracy (self-reported). Source: benchmarklist.com. Status: saturation imminent. 42 models tracked.

Top models

#ModelScore
1GPT-5 (Medium)47
2Gemini 2.5 Pro (Preview 06-05)46
3Gemini 2.5 Pro43.5
4O3 (Medium)37
5GPT-5 Mini (Medium)30.5
6Gemini 2.5 Flash23.5
7GPT-OSS-120B23.5
8DeepSeek V3.115.5
9O4 Mini (Medium)15.5
10DeepSeek R1 05289.5
11O3 Mini (Medium)9.5
12Grok 48
13O1 (Medium)8
14DeepSeek V3 (0324)7
15Qwen 3 235B A22B6

Interactive version: theaggregate.ai/benchmark?slug=ineqmath · How the rankings work · Data refreshed daily, snapshot 2026-07-22.