Skip to content

The Aggregate

Unified LLM rankings, daily

Models What's New Trends Benchmarks Skill Maps
How It Works Guesswork Metrics
The Aggregate
Loading data...

NatureBench — leaderboard

Metric: Surpass-SOTA (%). Source: frontisai.github.io. 12 models tracked.

Top models

#ModelScore
1Claude Opus 4.717.78
2Gemini 3.5 Flash15.56
3GLM-5.215.56
4GPT-5.514.44
5Claude Opus 4.612.22
6MiniMax-M311.11
7Qwen 3.7 Max10
8GPT-5.48.89
9Kimi K2.68.89
10GLM-5.17.78
11DeepSeek V4 Pro4.44
12MiniMax-M2.71.11

Interactive version: theaggregate.ai/benchmark?slug=naturebench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.

Built by Mikhail Doroshenko — AI researcher, co-author of Humanity’s Last Exam. Independent project; no affiliation with any AI lab.