Guesswork — LLMs vs a statistical aggregate at predicting benchmark scores

Guesswork is a live benchmark for predicting benchmark scores. Every day new benchmark results land, and this site has already predicted each one. It scores who called it closer: The Aggregate (the statistical model behind the unified rankings) or frontier LLMs given the same information. Predictions are made blind, before the score exists, and graded once it lands.

Current month 2026-09 is in progress: 10 of 50 shared predictions so far. Standings are published at 50.

Latest ranked month: 2026-08 · 103 shared predictions · 34 contestants. Ranked by mean absolute error in per-benchmark z-score units (MAE z, lower is better); “beats Aggregate” is the share of predictions where the contestant landed closer than this site's own.

2026-08 standings

#ContestantMAE (z)Beats Aggregate
1The Aggregate0.567
2Claude Opus 50.88749%
3Gemini 3.1 Pro Preview0.95935%
4Gemini 3.7 Flash0.96651%
5Grok 4.60.97543%
6Claude Sonnet 50.98151%
7Claude Opus 4.80.99747%
8GPT-5.51.00046%
9Ministral 3 14B1.01531%
10Kimi K2.61.02843%
11Gemma 3 4B1.05121%
12GPT-5.6 Luna1.05246%
13GPT-5.6 Terra1.06440%
14Mistral Small 31.07643%
15GPT-5.11.09839%
16Qwen 3.7 Max1.10550%
17GLM-5.21.12447%
18Grok 4.31.16543%
19Ling 2.6 Flash1.19632%
20Mistral Medium 3.51.20734%
21Codestral (2508)1.21234%
22Mistral Small 41.24733%
23Voxtral Small1.27731%
24Gemini 2.5 Pro1.28335%
25DeepSeek V4 Pro1.29240%

Earlier months: 2026-07, 2026-06.

Interactive version: theaggregate.ai/guesswork · How It Works · Data refreshed daily, snapshot 2026-09-05.