Our Benchmarks — evaluations The Aggregate runs itself
Everything else on this site is collected from public leaderboards. These are the evaluations The Aggregate runs itself: question banks it built, and Guesswork, the live contest for predicting benchmark scores. Apart from Guesswork, these results are shown on this page only and do not feed the rating.
Long-tail book recall
Closed-book questions about details in 60 public-domain books, from the famous to the obscure. Run Sep 2–9, 2026, 282 items.
Each question asks for one specific detail from a Project Gutenberg book, such as a minor character's name or the place where a scene happens, and the answer key is a verbatim span of the book. The 60 books sit on five rungs by how often their Wikipedia article is read, from the canon down to books almost nobody opens, so the bank keeps separating models long after the famous titles are solved.
Answers are matched against the key; an answer that misses the exact wording goes to a second-pass judge that accepts equivalent answers. Saying “unknown” counts as wrong. Reasoning effort was set to high.
| # | Model | Accuracy |
|---|---|---|
| 1 | Claude Fable 5.1 | 69.5% |
| 2 | Claude Fable 5 | 69.1% |
| 3 | Claude Opus 5 | 62.8% |
| 4 | Gemini 3.7 Flash | 56.7% |
| 5 | GPT-5.6 Sol | 55.3% |
| 6 | GPT-6 Astra | 50.7% |
| 7 | GPT-5.4 | 46.1% |
| 8 | DeepSeek V4 Pro | 44.0% |
| 9 | Claude Opus 4.8 | 42.2% |
| 10 | Muse Spark 1.3 | 39.4% |
| 11 | GPT-5.4 Mini | 38.3% |
| 12 | GPT-5.2 | 37.2% |
| 13 | Qwen 3.6 Plus | 31.9% |
| 14 | Kimi K2.6 | 30.5% |
| 15 | Kimi K2 0905 | 28.7% |
| 16 | Qwen 3.6 35B A3B | 27.7% |
| 17 | Claude Haiku 4.5 | 17.4% |
| 18 | GPT-OSS-120B | 17.4% |
Exotic fact recall
Closed-book lookups from five public registries: proteins, ISBNs, museum objects, asteroids and network ports. Run Sep 8–9, 2026, 85 items.
Each question asks for one exact field from a public registry: a UniProt protein entry, the title behind an ISBN, the accession number of a painting at the Met, the number of a named asteroid, or the TCP port that IANA assigns to a service. Every key is exact, so no judge is involved.
Each registry contributes well-known entries and a random sample from its long tail. The score counts the 85 questions that every model in the original panel answered, 49 of them from the tail. Reasoning effort was set to high.
| # | Model | Accuracy |
|---|---|---|
| 1 | GPT-5.6 Sol | 63.5% |
| 2 | Claude Fable 5.1 | 62.4% |
| 3 | Claude Fable 5 | 61.2% |
| 4 | GPT-6 Astra | 56.5% |
| 5 | Gemini 3.7 Flash | 55.3% |
| 6 | Claude Opus 5 | 50.6% |
| 7 | DeepSeek V4 Pro | 49.4% |
| 8 | GPT-5.4 | 41.2% |
| 9 | GPT-5.4 Mini | 36.5% |
| 10 | Qwen 3.6 Plus | 31.8% |
| 11 | Qwen 3.6 35B A3B | 27.1% |
| 12 | Claude Haiku 4.5 | 23.5% |
Asteroid numbers
Name 48 freshly sampled asteroids' catalogue numbers, from memory. Run Sep 9, 2026, 48 items.
A follow-up to the exotic recall bank, where DeepSeek and Gemini beat the Claude models on asteroids. 48 named asteroids were drawn at random from NASA/JPL's catalogue (numbers 1,000 to 20,000, first observed before 2000), and each model gave the permanent number for each name, closed book, graded by exact match. The reversal held on the fresh sample.
A reply cut off by the output limit counts as a miss; that happened at most once per model. Reasoning effort was set to high.
| # | Model | Accuracy |
|---|---|---|
| 1 | DeepSeek V4 Pro | 58.3% |
| 2 | Gemini 3.7 Flash | 56.2% |
| 3 | GPT-5.6 Sol | 52.1% |
| 4 | GPT-6 Astra | 47.9% |
| 5 | Claude Fable 5.1 | 33.3% |
| 6 | Claude Fable 5 | 27.1% |
| 7 | Claude Opus 5 | 20.8% |
Day-one recall check
The 34-question recall test every new free model gets on the day it appears. Run Aug 31 – Oct 3, 2026, 34 items.
A short version of the book recall bank: 34 questions about 12 public-domain books in three tiers, famous, mid and obscure. It is the first step of the battery we run on every new free model the day it appears, so it is the one recall board where free models sit next to the flagships.
Graded by matching the key, with no second-pass judge. Saying “unknown” counts as wrong; an empty reply does not count at all. Reasoning effort was set to high where the route accepts it.
| # | Model | Accuracy |
|---|---|---|
| 1 | Claude Fable 5 | 61.8% |
| 2 | Claude Opus 5 | 58.8% |
| 3 | Claude Fable 5.1 | 55.9% |
| 4 | GPT-5.5 | 47.1% |
| 5 | Claude Opus 4.8 | 41.2% |
| 6 | GPT-5.2 | 38.2% |
| 7 | GPT-5.4 | 38.2% |
| 8 | Kimi K2.6 | 32.4% |
| 9 | DeepSeek V4 Pro | 32.4% |
| 10 | Nex N2.5 Mini (free route) | 32.4% |
| 11 | Nemotron 3 Ultra (free route) | 32.3% |
| 12 | GPT-5.4 Mini | 29.4% |
| 13 | Qwen 3.6 35B A3B | 29.4% |
| 14 | Nex N2.5 Pro (free route) | 29.4% |
| 15 | Nemotron 3 Super (free route) | 26.5% |
| 16 | Kimi K2 0905 | 23.5% |
| 17 | Qwen 3.6 Plus | 23.5% |
| 18 | Ling-3.0-flash-fin (free route) | 23.5% |
| 19 | GPT-OSS-120B | 20.6% |
| 20 | Nemotron 3.5 Lightning (free route) | 17.6% |
| 21 | Ling-3.0-flash-sante (free route) | 17.6% |
| 22 | Ling-3.0-flash-VL (free route) | 17.6% |
| 23 | Apodex 1.1 Mini (free route) | 17.6% |
| 24 | Gemma 4 26B A4B (IT) (free route) | 14.7% |
| 25 | Gemma 4 31B (IT) (free route) | 14.7% |
| 26 | Space Bunny Alpha (stealth release) | 14.7% |
| 27 | Ling 3.1 Flash | 14.7% |
| 28 | Dots3-Note (Preview) (free route) | 11.8% |
| 29 | Claude Haiku 4.5 | 8.8% |
| 30 | Laguna XS 2.1 (free route) | 8.8% |
| 31 | North Mini Code (free route) | 8.8% |
| 32 | Laguna S 2.1 (free route) | 6.2% |
| 33 | LFM2.5-2.6B (free route) | 2.9% |
Forecasting resolved questions
Probabilities for 76 prediction-market questions that no model in the panel could look up from memory. Run Sep 1, 2026, 76 items.
Each question is a yes-or-no market on Manifold that resolved between May and August 2026. Before forecasting, every model is asked whether it already knows the outcome, and any question that any panel model answers correctly is dropped for everyone. That leaves 76 questions the models have to forecast rather than remember.
The score is the Brier score: the mean squared gap between the forecast probability and what happened, lower is better. Always answering 50% scores 0.250, and always answering the share of questions that resolved yes scores 0.249. No model beat that base rate: they were directionally right a little more often than chance and far too confident.
| # | Model | Brier score (lower is better) |
|---|---|---|
| Always the base rate | 0.249 | |
| Always 50% | 0.250 | |
| 1 | Claude Opus 5 | 0.273 |
| 2 | GPT-5.4 | 0.293 |
| 3 | Claude Opus 4.8 | 0.301 |
| 4 | GPT-5.5 | 0.305 |
| 5 | Claude Fable 5 | 0.311 |
| 6 | GPT-5.2 | 0.328 |
| 7 | DeepSeek V4 Pro | 0.330 |
| 8 | Kimi K2.6 | 0.335 |
| 9 | Qwen 3.6 Plus | 0.355 |
Inversion and execution ladder
24 generated puzzles: trace a short program, or undo a hidden digit transform, at rising difficulty. Run Aug 28 – Oct 3, 2026, 24 items.
Half the items give the model a short generated program and ask for its output; the other half give the result of a hidden digit transform and ask for the input that produced it. Six rungs of difficulty and exact-match answers, on items generated for this ladder from a fixed seed.
Run at each provider's default reasoning effort, which this ladder is very sensitive to. The scores show what a model does with its out-of-the-box settings, not its ceiling.
| # | Model | Accuracy |
|---|---|---|
| 1 | GPT-5.5 | 100.0% |
| 2 | Dots3-Note (Preview) (free route) | 100.0% |
| 3 | GPT-5.2 | 91.7% |
| 4 | GLM-5.3 (Mistral API) | 91.7% |
| 5 | Ling 3.1 Flash | 91.7% |
| 6 | Kimi K2.6 | 87.5% |
| 7 | Claude Opus 4.8 | 83.3% |
| 8 | Nemotron 3 Super (free route) | 83.3% |
| 9 | GPT-OSS-120B | 79.2% |
| 10 | Ling-3.0-flash-VL (free route) | 78.3% |
| 11 | Nex N2.5 Pro (free route) | 70.8% |
| 12 | Nex N2.5 Mini (free route) | 50.0% |
| 13 | GPT-5.4 | 41.7% |
| 14 | Gemma 4 26B A4B (IT) (free route) | 41.7% |
| 15 | Space Bunny Alpha (stealth release) | 26.1% |
| 16 | Gemma 4 31B (IT) (free route) | 25.0% |
| 17 | Kimi K2 0905 | 20.8% |
Puzzle portfolio ladder
30 generated puzzles in five families: logic grids, knapsacks, document facts, specifications and bug hunts. Run Aug 29 – Oct 3, 2026, 30 items.
Six items from each of five generators whose answers are fixed by construction: zebra logic grids, knapsack optimisation, fact lookup in a long document, checking an implementation against its specification, and finding a planted bug. Difficulty rises within each family, and the bank is frozen.
Run at each provider's default reasoning effort, like the ladder above.
| # | Model | Accuracy |
|---|---|---|
| 1 | GPT-5.2 | 96.7% |
| 2 | GPT-5.5 | 96.7% |
| 3 | Ling 3.1 Flash | 96.3% |
| 4 | GLM-5.3 (Mistral API) | 92.9% |
| 5 | Kimi K2.6 | 90.0% |
| 6 | Claude Opus 4.8 | 90.0% |
| 7 | Qwen 3.6 Plus | 86.7% |
| 8 | DeepSeek V4 Pro | 86.7% |
| 9 | GPT-OSS-120B | 83.3% |
| 10 | Nex N2.5 Pro (free route) | 82.1% |
| 11 | Nemotron 3 Ultra (free route) | 77.8% |
| 12 | Gemma 4 31B (IT) (free route) | 73.3% |
| 13 | Nemotron 3 Super (free route) | 66.7% |
| 14 | Ling-3.0-flash-VL (free route) | 63.3% |
| 15 | GPT-5.4 | 60.0% |
| 16 | Space Bunny Alpha (stealth release) | 60.0% |
| 17 | Gemma 4 26B A4B (IT) (free route) | 53.3% |
| 18 | Laguna XS 2.1 (free route) | 44.8% |
| 19 | Kimi K2 0905 | 43.3% |
Also run here
- MEGA-Bench Hanoi re-run. Gemini 3.8 Flash on MEGA-Bench's 14-item Tower of Hanoi planning task, scored 58.2% with MEGA-Bench's own prompts and scorer (September 20). The row sits on the public board, and the check found five of its fourteen reference answers impossible.
- Predicting a small model from its name. Gemini 3.8 Flash predicted, question by question, where Llama 3.2 1B would fail on 498 MMLU-Pro and GSM8K items, knowing only its name: AUROC 0.71, but it took the 33% model for a 66% one (September 21).
- Verbatim recital. Continuing a Project Gutenberg paragraph or an Apollo 11 transcript window word for word from a 12-word cue: every model tested scored 0 of 30 paragraphs and 0 of 24 transcript windows at 100 words (September 9).
- Classifier qualification suite. The tests a model must pass before it becomes the benchmark classifier behind Task Explorer.
Interactive version: theaggregate.ai/our-benchmarks · How It Works · Data refreshed daily, snapshot 2026-10-09.