Guesswork — LLMs vs a statistical aggregate at predicting benchmark scores

Guesswork is a live benchmark for predicting benchmark scores. Every day new benchmark results land, and this site has already predicted each one. It scores who called it closer: The Aggregate (the statistical model behind the unified rankings) or frontier LLMs given the same information. Predictions are made blind, before the score exists, and graded once it lands.

Latest ranked month: 2026-09 · 54 shared predictions · 33 contestants. Ranked by mean absolute error in per-benchmark z-score units (MAE z, lower is better); “beats Aggregate” is the share of predictions where the contestant landed closer than this site's own.

2026-09 standings

#ContestantMAE (z)Beats Aggregate
1The Aggregate0.407
2Qwen 3.7 Max0.58843%
3Claude Sonnet 50.61241%
4Claude Opus 50.62137%
5Claude Opus 4.80.62437%
6Grok 4.60.63039%
7GPT-5.10.64437%
8GPT-5.50.64941%
9GPT-5.6 Terra0.65835%
10Grok 4.30.65933%
11GPT-5.6 Luna0.68639%
12GLM-5.20.68739%
13Gemini 3.7 Flash0.70435%
14Kimi K2.60.73041%
15DeepSeek V4 Pro0.81939%
16Mistral Small 30.83335%
17Gemini 2.5 Pro0.86235%
18Gemini 3.1 Pro Preview0.87332%
19Mistral Medium 3.50.90833%
20Mistral Small 40.93328%
21Voxtral Small0.95528%
22Ministral 3 8B0.96526%
23Codestral (2508)0.96832%
24Mistral Large 31.02724%
25GPT-OSS 20B1.03124%

Earlier months: 2026-08, 2026-07, 2026-06.

Interactive version: theaggregate.ai/guesswork · How It Works · Data refreshed daily, snapshot 2026-09-19.