Guesswork — LLMs vs a statistical aggregate at predicting benchmark scores
Guesswork is a live benchmark for predicting benchmark scores. Every day new benchmark results land, and this site has already predicted each one. It scores who called it closer: The Aggregate (the statistical model behind the unified rankings) or frontier LLMs given the same information. Predictions are made blind, before the score exists, and graded once it lands.
Current month 2026-09 is in progress: 10 of 50 shared predictions so far. Standings are published at 50.
Latest ranked month: 2026-08 · 103 shared predictions · 34 contestants. Ranked by mean absolute error in per-benchmark z-score units (MAE z, lower is better); “beats Aggregate” is the share of predictions where the contestant landed closer than this site's own.
2026-08 standings
| # | Contestant | MAE (z) | Beats Aggregate |
|---|---|---|---|
| 1 | The Aggregate | 0.567 | — |
| 2 | Claude Opus 5 | 0.887 | 49% |
| 3 | Gemini 3.1 Pro Preview | 0.959 | 35% |
| 4 | Gemini 3.7 Flash | 0.966 | 51% |
| 5 | Grok 4.6 | 0.975 | 43% |
| 6 | Claude Sonnet 5 | 0.981 | 51% |
| 7 | Claude Opus 4.8 | 0.997 | 47% |
| 8 | GPT-5.5 | 1.000 | 46% |
| 9 | Ministral 3 14B | 1.015 | 31% |
| 10 | Kimi K2.6 | 1.028 | 43% |
| 11 | Gemma 3 4B | 1.051 | 21% |
| 12 | GPT-5.6 Luna | 1.052 | 46% |
| 13 | GPT-5.6 Terra | 1.064 | 40% |
| 14 | Mistral Small 3 | 1.076 | 43% |
| 15 | GPT-5.1 | 1.098 | 39% |
| 16 | Qwen 3.7 Max | 1.105 | 50% |
| 17 | GLM-5.2 | 1.124 | 47% |
| 18 | Grok 4.3 | 1.165 | 43% |
| 19 | Ling 2.6 Flash | 1.196 | 32% |
| 20 | Mistral Medium 3.5 | 1.207 | 34% |
| 21 | Codestral (2508) | 1.212 | 34% |
| 22 | Mistral Small 4 | 1.247 | 33% |
| 23 | Voxtral Small | 1.277 | 31% |
| 24 | Gemini 2.5 Pro | 1.283 | 35% |
| 25 | DeepSeek V4 Pro | 1.292 | 40% |
Earlier months: 2026-07, 2026-06.
Interactive version: theaggregate.ai/guesswork · How It Works · Data refreshed daily, snapshot 2026-09-05.