BenchBench — leaderboard

Meta-benchmark leaderboard aggregating how well models perform across benchmark-oriented evaluation tasks.

Metric: Aggregate Score (%). Source: huggingface.co. Status: saturated. 137 models tracked.

Top models

#ModelScore
1GPT-4o (2024-05-13)97.67
2GPT-4o ChatGPT97.54
3GPT-4o (2024-08-06)96.53
4Claude 3.5 Sonnet (20240620)95.73
5Gemini 1.5 Pro (Preview 0801)95.45
6Llama 3.1 70B Instruct93.43
7GPT-4 Turbo90.56
8Claude 3 Opus (20240229)88.24
9Yi Large (Preview)87.14
10Llama 3.1 405B Instruct85.98
11GPT-4 Preview (0125)84.92
12Hermes 3 - Llama-3.1 70B84.51
13zephyr-orpo-141B-A35B-v0.184.14
14Mistral Large 2 (Jul)83.75
15GPT-4o Mini (2024-07-18)83.49

Interactive version: theaggregate.ai/benchmark?slug=benchbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.