Guesswork — leaderboard

This site's own live benchmark for predicting benchmark scores: frontier and budget LLMs forecast newly published (model, benchmark) results a day before they land, scored in per-benchmark z-units (MAE, lower is better) against the same-information statistical baseline The Aggregate.

Metric: MAE (z-score units). Source: aibenchmarks.dev. Status: saturated. 21 models tracked.

Top models

#ModelScore
1GLM-5.20.66
2Gemini 3.1 Pro (Preview)0.76
3Gemini 2.5 Pro0.8
4Claude Sonnet 50.8
5DeepSeek V4 Pro0.8
6Qwen 3.7 Max0.85
7Grok 4.30.89
8Kimi K2.60.96
9GPT-5.50.96
10Claude Opus 4.80.98
11GPT-5.6 Terra1.01
12Ling-2.6-flash1.01
13GPT-5.11.02
14Qwen 3.5 9B1.02
15GPT-OSS-20B1.03

Interactive version: theaggregate.ai/benchmark?slug=guesswork · How the rankings work · Data refreshed daily, snapshot 2026-07-22.