BenchBench — leaderboard
Meta-benchmark leaderboard aggregating how well models perform across benchmark-oriented evaluation tasks.
Metric: Aggregate Score (%). Source: huggingface.co. Status: saturated. 137 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4o (2024-05-13) | 97.67 |
| 2 | GPT-4o ChatGPT | 97.54 |
| 3 | GPT-4o (2024-08-06) | 96.53 |
| 4 | Claude 3.5 Sonnet (20240620) | 95.73 |
| 5 | Gemini 1.5 Pro (Preview 0801) | 95.45 |
| 6 | Llama 3.1 70B Instruct | 93.43 |
| 7 | GPT-4 Turbo | 90.56 |
| 8 | Claude 3 Opus (20240229) | 88.24 |
| 9 | Yi Large (Preview) | 87.14 |
| 10 | Llama 3.1 405B Instruct | 85.98 |
| 11 | GPT-4 Preview (0125) | 84.92 |
| 12 | Hermes 3 - Llama-3.1 70B | 84.51 |
| 13 | zephyr-orpo-141B-A35B-v0.1 | 84.14 |
| 14 | Mistral Large 2 (Jul) | 83.75 |
| 15 | GPT-4o Mini (2024-07-18) | 83.49 |
Interactive version: theaggregate.ai/benchmark?slug=benchbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.