The Aggregate: unified LLM rankings, daily
We fold every public benchmark we can find into one daily ranking. Unified rankings for 1761 AI models across 5615 public benchmarks, aggregated from public leaderboards and refreshed daily. Rankings use an Item Response Theory (IRT) fit that maps every model onto one ELO scale.
Top models by unified ELO
| Rank | Model | Provider | Unified ELO | Benchmarks |
|---|---|---|---|---|
| 1 | Claude Fable 5.1 | Anthropic | 1808 ± 1 | 157 |
| 2 | GPT-6 | OpenAI | 1805 ± 1 | 100 |
| 3 | Claude Mythos 5 | Anthropic | 1800 ± 1 | 54 |
| 4 | Gemini 3.8 Flash | 1778 ± 1 | 146 | |
| 5 | Gemini 3.7 Flash | 1774 ± 1 | 267 | |
| 6 | Claude Fable 5 | Anthropic | 1773 ± 1 | 412 |
| 7 | GPT-5.5 Pro | OpenAI | 1773 ± 1 | 55 |
| 8 | GPT-5.6 Pro Sol | OpenAI | 1772 ± 1 | 18 |
| 9 | Claude Opus 5 | Anthropic | 1771 ± 1 | 393 |
| 10 | Muse Spark 1.3 | Meta | 1770 ± 1 | 83 |
| 11 | Claude Mythos Preview | Anthropic | 1767 ± 1 | 79 |
| 12 | GPT-5.6 Sol | OpenAI | 1761 ± 1 | 516 |
| 13 | Kimi K3 | Moonshot | 1754 ± 1 | 403 |
| 14 | GPT-5.4 Pro | OpenAI | 1753 ± 1 | 62 |
| 15 | GPT-5.5 | OpenAI | 1743 ± 1 | 1080 |
| 16 | Gemini 3 Deep Think | 1742 ± 1 | 21 | |
| 17 | Grok 4.6 | xAI | 1741 ± 1 | 226 |
| 18 | Gemini 3.6 Flash | 1737 ± 1 | 299 | |
| 19 | Muse Spark 1.2 | Meta | 1737 ± 1 | 161 |
| 20 | Qwen 3.8 Max | Alibaba | 1735 ± 1 | 228 |
Latest leaderboard changes
- GPT-6 took the lead on RuneBench (34063 vs 12757 by Grok 4.6)
- GPT-5.5 (xHigh) took the lead on LiveBench Zebra Puzzle (100 vs 48 by Claude 3.5 Sonnet (20240620))
- Claude Opus 5 (Max) took the lead on LiveBench AMPS Hard (99.01 vs 49 by Claude 3.5 Sonnet (20240620))
- Claude Opus 4.7 (xHigh) took the lead on LiveBench Math Comp (98.04 vs 54.17 by Gemini 1.5 Pro (Preview 0827))
- Gemini 3.1 Pro (Preview) (High) took the lead on LiveBench Connections (100 vs 58 by GPT-4o (2024-08-06))
- GPT-6 (Max) took the lead on LiveBench Plot Unscrambling (86.29 vs 55.14 by Claude 3.5 Sonnet (20240620))
- DeepSeek V4 Pro took the lead on LiveBench Table Reformat (100 vs 70 by GPT-4 Preview (0125))
- Claude Fable 5 (Max) took the lead on LiveBench Typos (94 vs 68 by Claude 3 Opus (20240229))
Explore: all models · benchmark catalog · capability trends · daily changes
Interactive version: theaggregate.ai/ · How It Works · Data refreshed daily, snapshot 2026-09-05.