MathNet - Geometry: leaderboard

Metric: Problem-solving accuracy (%) on the geometry problems on MathNet-Solve-Test, Olympiad problems from official national competition booklets of 47 countries; a GPT-5 grader scores each solution 0-7 against the official solution and 6 or more counts as correct; image-capable models get the figures, text-only models a text description of them; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 14 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)74.6
2Gemini 3 Flash (Preview)67
3GPT-561.1
4GPT-5 Mini50.3
5Claude Opus 4.6 (Thinking)44.3
6Gemini 2.5 Flash36.8
7GPT-5 Nano32.4
8DeepSeek V3.2 (Non-reasoning)32.2
9Grok 321.7
10GPT-4.115.7
11Llama 4 Maverick Instruct FP810.7
12GPT-4o4.5
13Ministral 3B4.3

Interactive version: theaggregate.ai/benchmark?slug=mathnet-geometry · How It Works · Data refreshed daily, snapshot 2026-10-07.