Skip to content

The Aggregate

Unified LLM rankings, daily

Models What's New Trends Benchmarks Skill Maps
How It Works Guesswork Metrics
The Aggregate
Loading data...

SmellBench — leaderboard

Metric: Weighted Effectiveness (E) (self-reported). Source: benchmarklist.com. 9 models tracked.

Top models

#ModelScore
1GPT-5.3 Codex47.8
2Gemini 3.1 Pro (Preview)43.5
3GPT-5.442.8
4Claude Haiku 4.537.9
5Claude Opus 4.637
6Claude Sonnet 4.635.4

Interactive version: theaggregate.ai/benchmark?slug=smellbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.

Built by Mikhail Doroshenko — AI researcher, co-author of Humanity’s Last Exam. Independent project; no affiliation with any AI lab.