Skip to content

The Aggregate

Unified LLM rankings, daily

Models What's New Trends Benchmarks Skill Maps
How It Works Guesswork Metrics
The Aggregate
Loading data...

HELM Classic - BLiMP — leaderboard

Metric: Exact Match (%). Source: crfm.stanford.edu. 32 models tracked.

Top models

#ModelScore
1davinci83.97
2gpt-neox-20B83.85
3gpt-j-6B83.4
4text-davinci-00382.25
5text-davinci-00281.2
6text-babbage-00176.95
7text-curie-00176.05
8text-ada-00172.78

Interactive version: theaggregate.ai/benchmark?slug=helm-classic-blimp · How the rankings work · Data refreshed daily, snapshot 2026-07-22.

Built by Mikhail Doroshenko — AI researcher, co-author of Humanity’s Last Exam. Independent project; no affiliation with any AI lab.