Skip to content

The Aggregate

Unified LLM rankings, daily

Models What's New Trends Benchmarks Skill Maps
How It Works Guesswork Metrics
The Aggregate
Loading data...

LLM Stats (Internal API instruction following (hard)) — leaderboard

Metric: Score (%). Source: llm-stats.com. 7 models tracked.

Top models

#ModelScore
1GPT-564
2GPT-4.554
3O3 Mini50
4GPT-4.149.1
5GPT-4.1 Mini45.1
6GPT-4.1 Nano31.6
7GPT-4o29.2

Interactive version: theaggregate.ai/benchmark?slug=llm-stats-internal-api-instruction-following-hard · How the rankings work · Data refreshed daily, snapshot 2026-07-22.

Built by Mikhail Doroshenko — AI researcher, co-author of Humanity’s Last Exam. Independent project; no affiliation with any AI lab.