BenchBench: leaderboard

Meta-benchmark leaderboard aggregating how well models perform across benchmark-oriented evaluation tasks.

Metric: Aggregate Score (%). Source: huggingface.co. Status: saturated. 137 models tracked.

Top models

#ModelScore
1GPT-4o (2024-05-13)97.67
2GPT-4o ChatGPT97.54
3GPT-4o (2024-08-06)96.53
4Claude 3.5 Sonnet (20240620)95.73
5Llama 3.1 70B Instruct93.43
6GPT-4 Turbo90.56
7Claude 3 Opus (20240229)88.24
8Yi Large (Preview)87.14
9Llama 3.1 405B Instruct85.98
10GPT-4 Preview (0125)84.92
11Hermes 3 - Llama-3.1 70B84.51
12zephyr-orpo-141B-A35B-v0.184.14
13Mistral Large 2 (Jul)83.75
14GPT-4o Mini (2024-07-18)83.49
15Claude 283.33

Interactive version: theaggregate.ai/benchmark?slug=benchbench · How It Works · Data refreshed daily, snapshot 2026-09-05.