How It Works

Every day the pipeline scrapes hundreds of public leaderboards, normalizes model names, and fuses duplicate sources into one score matrix. A 1-parameter IRT (Rasch) model is then fit over the matrix: each model gets an ability estimate and each benchmark a difficulty estimate. Abilities map to an ELO scale centered at 1500.

Benchmarks that anticorrelate with general ability, have very low g-loading, or are too sparse for statistically meaningful ranking are excluded from the IRT fit but stay visible in the catalog. A hybrid model combining ability, skill axes, and provider factors predicts missing scores, and saturation forecasts estimate when benchmarks stop discriminating between frontier models. Prediction accuracy is tracked against later scraped results.

The Aggregate is built and run by Mikhail Doroshenko, an AI researcher and co-author of Humanity’s Last Exam. It is an independent project with no affiliation with any AI lab. Every ranked score comes from a public leaderboard or a published report and links back to where it came from; the evaluations the site runs itself, such as Guesswork, are shown as its own work.

The changelog on What’s New is computed by the pipeline with no language model involved. The story above it is written by Claude from that changelog during the morning audit. When a published number turns out to be wrong, the fix ships with the next daily publish and the boards it touched appear in What’s New as changes; corrections that changed what a visitor saw are listed on the interactive page.

Interactive version: theaggregate.ai/how-it-works · How It Works · Data refreshed daily, snapshot 2026-09-19.