Guesswork — LLMs vs a statistical aggregate at predicting benchmark scores
Guesswork is a live benchmark for predicting benchmark scores. Every day new benchmark results land, and this site has already predicted each one. It scores who called it closer: The Aggregate (the statistical model behind the unified rankings) or frontier LLMs given the same information. Predictions are made blind, before the score exists, and graded once it lands.
Latest ranked month: 2026-09 · 54 shared predictions · 33 contestants. Ranked by mean absolute error in per-benchmark z-score units (MAE z, lower is better); “beats Aggregate” is the share of predictions where the contestant landed closer than this site's own.
2026-09 standings
| # | Contestant | MAE (z) | Beats Aggregate |
|---|---|---|---|
| 1 | The Aggregate | 0.407 | — |
| 2 | Qwen 3.7 Max | 0.588 | 43% |
| 3 | Claude Sonnet 5 | 0.612 | 41% |
| 4 | Claude Opus 5 | 0.621 | 37% |
| 5 | Claude Opus 4.8 | 0.624 | 37% |
| 6 | Grok 4.6 | 0.630 | 39% |
| 7 | GPT-5.1 | 0.644 | 37% |
| 8 | GPT-5.5 | 0.649 | 41% |
| 9 | GPT-5.6 Terra | 0.658 | 35% |
| 10 | Grok 4.3 | 0.659 | 33% |
| 11 | GPT-5.6 Luna | 0.686 | 39% |
| 12 | GLM-5.2 | 0.687 | 39% |
| 13 | Gemini 3.7 Flash | 0.704 | 35% |
| 14 | Kimi K2.6 | 0.730 | 41% |
| 15 | DeepSeek V4 Pro | 0.819 | 39% |
| 16 | Mistral Small 3 | 0.833 | 35% |
| 17 | Gemini 2.5 Pro | 0.862 | 35% |
| 18 | Gemini 3.1 Pro Preview | 0.873 | 32% |
| 19 | Mistral Medium 3.5 | 0.908 | 33% |
| 20 | Mistral Small 4 | 0.933 | 28% |
| 21 | Voxtral Small | 0.955 | 28% |
| 22 | Ministral 3 8B | 0.965 | 26% |
| 23 | Codestral (2508) | 0.968 | 32% |
| 24 | Mistral Large 3 | 1.027 | 24% |
| 25 | GPT-OSS 20B | 1.031 | 24% |
Earlier months: 2026-08, 2026-07, 2026-06.
Interactive version: theaggregate.ai/guesswork · How It Works · Data refreshed daily, snapshot 2026-09-19.