HiddenMath: leaderboard
DeepMind mathematical reasoning evaluation using hidden problems to reduce memorization and test whether models can solve unfamiliar contest-style tasks.
Metric: Score (%). Source: llm-stats.com. Status: saturated. 13 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.0 Flash | 63 |
| 2 | Gemma 3 27B | 60.3 |
| 3 | Gemini 2.0 Flash Lite | 55.3 |
| 4 | Gemma 3 12B | 54.5 |
| 5 | Gemini 1.5 Pro | 52 |
| 6 | Gemini 1.5 Flash | 47.2 |
| 7 | Gemma 3 4B | 43 |
| 8 | Gemma 3n E4B Instruct | 37.7 |
| 9 | Gemini 1.5 Flash-8B | 32.8 |
| 10 | gemma-3n-E2B-it Instruct | 27.7 |
| 11 | Gemma 3 1B | 15.8 |
Interactive version: theaggregate.ai/benchmark?slug=hiddenmath · How It Works · Data refreshed daily, snapshot 2026-09-05.