MathNet: leaderboard

Metric: Problem-solving accuracy (%) over all problems (micro average) on MathNet-Solve-Test, Olympiad problems from official national competition booklets of 47 countries; a GPT-5 grader scores each solution 0-7 against the official solution and 6 or more counts as correct; image-capable models get the figures, text-only models a text description of them; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 14 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)78.4
2Gemini 3 Flash (Preview)70.4
3GPT-569.3
4GPT-5 Mini57
5Claude Opus 4.6 (Thinking)45.7
6GPT-5 Nano42.2
7Gemini 2.5 Flash41.1
8DeepSeek V3.2 (Non-reasoning)40.1
9Grok 328.5
10GPT-4.121.4
11Llama 4 Maverick Instruct FP814.7
12GPT-4o6.8
13Ministral 3B4.4

Interactive version: theaggregate.ai/benchmark?slug=mathnet · How It Works · Data refreshed daily, snapshot 2026-10-07.