How the Rankings Work

Every day the pipeline scrapes hundreds of public leaderboards, normalizes model names, and fuses duplicate sources into one score matrix. A 1-parameter IRT (Rasch) model is then fit over the matrix: each model gets an ability estimate and each benchmark a difficulty estimate. Abilities map to an ELO scale centered at 1500.

Benchmarks that anticorrelate with general ability, have very low g-loading, or are too sparse for statistically meaningful ranking are excluded from the IRT fit but stay visible in the catalog. A hybrid model combining ability, skill axes, and provider factors predicts missing scores, and saturation forecasts estimate when benchmarks stop discriminating between frontier models. Prediction accuracy is tracked against later scraped results.

Interactive version: theaggregate.ai/methodology · How the rankings work · Data refreshed daily, snapshot 2026-07-22.