Our Benchmarks — evaluations The Aggregate runs itself

Everything else on this site is collected from public leaderboards. These are the evaluations The Aggregate runs itself: question banks it built, and Guesswork, the live contest for predicting benchmark scores. Apart from Guesswork, these results are shown on this page only and do not feed the rating.

Long-tail book recall

Closed-book questions about details in 60 public-domain books, from the famous to the obscure. Run Sep 2–9, 2026, 282 items.

Each question asks for one specific detail from a Project Gutenberg book, such as a minor character's name or the place where a scene happens, and the answer key is a verbatim span of the book. The 60 books sit on five rungs by how often their Wikipedia article is read, from the canon down to books almost nobody opens, so the bank keeps separating models long after the famous titles are solved.

Answers are matched against the key; an answer that misses the exact wording goes to a second-pass judge that accepts equivalent answers. Saying “unknown” counts as wrong. Reasoning effort was set to high.

#ModelAccuracy
1Claude Fable 5.169.5%
2Claude Fable 569.1%
3Claude Opus 562.8%
4Gemini 3.7 Flash56.7%
5GPT-5.6 Sol55.3%
6GPT-6 Astra50.7%
7GPT-5.446.1%
8DeepSeek V4 Pro44.0%
9Claude Opus 4.842.2%
10Muse Spark 1.339.4%
11GPT-5.4 Mini38.3%
12GPT-5.237.2%
13Qwen 3.6 Plus31.9%
14Kimi K2.630.5%
15Kimi K2 090528.7%
16Qwen 3.6 35B A3B27.7%
17Claude Haiku 4.517.4%
18GPT-OSS-120B17.4%

Exotic fact recall

Closed-book lookups from five public registries: proteins, ISBNs, museum objects, asteroids and network ports. Run Sep 8–9, 2026, 85 items.

Each question asks for one exact field from a public registry: a UniProt protein entry, the title behind an ISBN, the accession number of a painting at the Met, the number of a named asteroid, or the TCP port that IANA assigns to a service. Every key is exact, so no judge is involved.

Each registry contributes well-known entries and a random sample from its long tail. The score counts the 85 questions that every model in the original panel answered, 49 of them from the tail. Reasoning effort was set to high.

#ModelAccuracy
1GPT-5.6 Sol63.5%
2Claude Fable 5.162.4%
3Claude Fable 561.2%
4GPT-6 Astra56.5%
5Gemini 3.7 Flash55.3%
6Claude Opus 550.6%
7DeepSeek V4 Pro49.4%
8GPT-5.441.2%
9GPT-5.4 Mini36.5%
10Qwen 3.6 Plus31.8%
11Qwen 3.6 35B A3B27.1%
12Claude Haiku 4.523.5%

Asteroid numbers

Name 48 freshly sampled asteroids' catalogue numbers, from memory. Run Sep 9, 2026, 48 items.

A follow-up to the exotic recall bank, where DeepSeek and Gemini beat the Claude models on asteroids. 48 named asteroids were drawn at random from NASA/JPL's catalogue (numbers 1,000 to 20,000, first observed before 2000), and each model gave the permanent number for each name, closed book, graded by exact match. The reversal held on the fresh sample.

A reply cut off by the output limit counts as a miss; that happened at most once per model. Reasoning effort was set to high.

#ModelAccuracy
1DeepSeek V4 Pro58.3%
2Gemini 3.7 Flash56.2%
3GPT-5.6 Sol52.1%
4GPT-6 Astra47.9%
5Claude Fable 5.133.3%
6Claude Fable 527.1%
7Claude Opus 520.8%

Day-one recall check

The 34-question recall test every new free model gets on the day it appears. Run Aug 31 – Oct 3, 2026, 34 items.

A short version of the book recall bank: 34 questions about 12 public-domain books in three tiers, famous, mid and obscure. It is the first step of the battery we run on every new free model the day it appears, so it is the one recall board where free models sit next to the flagships.

Graded by matching the key, with no second-pass judge. Saying “unknown” counts as wrong; an empty reply does not count at all. Reasoning effort was set to high where the route accepts it.

#ModelAccuracy
1Claude Fable 561.8%
2Claude Opus 558.8%
3Claude Fable 5.155.9%
4GPT-5.547.1%
5Claude Opus 4.841.2%
6GPT-5.238.2%
7GPT-5.438.2%
8Kimi K2.632.4%
9DeepSeek V4 Pro32.4%
10Nex N2.5 Mini (free route)32.4%
11Nemotron 3 Ultra (free route)32.3%
12GPT-5.4 Mini29.4%
13Qwen 3.6 35B A3B29.4%
14Nex N2.5 Pro (free route)29.4%
15Nemotron 3 Super (free route)26.5%
16Kimi K2 090523.5%
17Qwen 3.6 Plus23.5%
18Ling-3.0-flash-fin (free route)23.5%
19GPT-OSS-120B20.6%
20Nemotron 3.5 Lightning (free route)17.6%
21Ling-3.0-flash-sante (free route)17.6%
22Ling-3.0-flash-VL (free route)17.6%
23Apodex 1.1 Mini (free route)17.6%
24Gemma 4 26B A4B (IT) (free route)14.7%
25Gemma 4 31B (IT) (free route)14.7%
26Space Bunny Alpha (stealth release)14.7%
27Ling 3.1 Flash14.7%
28Dots3-Note (Preview) (free route)11.8%
29Claude Haiku 4.58.8%
30Laguna XS 2.1 (free route)8.8%
31North Mini Code (free route)8.8%
32Laguna S 2.1 (free route)6.2%
33LFM2.5-2.6B (free route)2.9%

Forecasting resolved questions

Probabilities for 76 prediction-market questions that no model in the panel could look up from memory. Run Sep 1, 2026, 76 items.

Each question is a yes-or-no market on Manifold that resolved between May and August 2026. Before forecasting, every model is asked whether it already knows the outcome, and any question that any panel model answers correctly is dropped for everyone. That leaves 76 questions the models have to forecast rather than remember.

The score is the Brier score: the mean squared gap between the forecast probability and what happened, lower is better. Always answering 50% scores 0.250, and always answering the share of questions that resolved yes scores 0.249. No model beat that base rate: they were directionally right a little more often than chance and far too confident.

#ModelBrier score (lower is better)
Always the base rate0.249
Always 50%0.250
1Claude Opus 50.273
2GPT-5.40.293
3Claude Opus 4.80.301
4GPT-5.50.305
5Claude Fable 50.311
6GPT-5.20.328
7DeepSeek V4 Pro0.330
8Kimi K2.60.335
9Qwen 3.6 Plus0.355

Inversion and execution ladder

24 generated puzzles: trace a short program, or undo a hidden digit transform, at rising difficulty. Run Aug 28 – Oct 3, 2026, 24 items.

Half the items give the model a short generated program and ask for its output; the other half give the result of a hidden digit transform and ask for the input that produced it. Six rungs of difficulty and exact-match answers, on items generated for this ladder from a fixed seed.

Run at each provider's default reasoning effort, which this ladder is very sensitive to. The scores show what a model does with its out-of-the-box settings, not its ceiling.

#ModelAccuracy
1GPT-5.5100.0%
2Dots3-Note (Preview) (free route)100.0%
3GPT-5.291.7%
4GLM-5.3 (Mistral API)91.7%
5Ling 3.1 Flash91.7%
6Kimi K2.687.5%
7Claude Opus 4.883.3%
8Nemotron 3 Super (free route)83.3%
9GPT-OSS-120B79.2%
10Ling-3.0-flash-VL (free route)78.3%
11Nex N2.5 Pro (free route)70.8%
12Nex N2.5 Mini (free route)50.0%
13GPT-5.441.7%
14Gemma 4 26B A4B (IT) (free route)41.7%
15Space Bunny Alpha (stealth release)26.1%
16Gemma 4 31B (IT) (free route)25.0%
17Kimi K2 090520.8%

Puzzle portfolio ladder

30 generated puzzles in five families: logic grids, knapsacks, document facts, specifications and bug hunts. Run Aug 29 – Oct 3, 2026, 30 items.

Six items from each of five generators whose answers are fixed by construction: zebra logic grids, knapsack optimisation, fact lookup in a long document, checking an implementation against its specification, and finding a planted bug. Difficulty rises within each family, and the bank is frozen.

Run at each provider's default reasoning effort, like the ladder above.

#ModelAccuracy
1GPT-5.296.7%
2GPT-5.596.7%
3Ling 3.1 Flash96.3%
4GLM-5.3 (Mistral API)92.9%
5Kimi K2.690.0%
6Claude Opus 4.890.0%
7Qwen 3.6 Plus86.7%
8DeepSeek V4 Pro86.7%
9GPT-OSS-120B83.3%
10Nex N2.5 Pro (free route)82.1%
11Nemotron 3 Ultra (free route)77.8%
12Gemma 4 31B (IT) (free route)73.3%
13Nemotron 3 Super (free route)66.7%
14Ling-3.0-flash-VL (free route)63.3%
15GPT-5.460.0%
16Space Bunny Alpha (stealth release)60.0%
17Gemma 4 26B A4B (IT) (free route)53.3%
18Laguna XS 2.1 (free route)44.8%
19Kimi K2 090543.3%

Also run here

  • MEGA-Bench Hanoi re-run. Gemini 3.8 Flash on MEGA-Bench's 14-item Tower of Hanoi planning task, scored 58.2% with MEGA-Bench's own prompts and scorer (September 20). The row sits on the public board, and the check found five of its fourteen reference answers impossible.
  • Predicting a small model from its name. Gemini 3.8 Flash predicted, question by question, where Llama 3.2 1B would fail on 498 MMLU-Pro and GSM8K items, knowing only its name: AUROC 0.71, but it took the 33% model for a 66% one (September 21).
  • Verbatim recital. Continuing a Project Gutenberg paragraph or an Apollo 11 transcript window word for word from a 12-word cue: every model tested scored 0 of 30 paragraphs and 0 of 24 transcript windows at 100 words (September 9).
  • Classifier qualification suite. The tests a model must pass before it becomes the benchmark classifier behind Task Explorer.

Interactive version: theaggregate.ai/our-benchmarks · How It Works · Data refreshed daily, snapshot 2026-10-09.