BBH — leaderboard

Big-Bench Hard benchmark with challenging tasks requiring multi-step reasoning.

Metric: Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 16 models tracked.

Top models

#ModelScore
1DeepSeek V387.5
2Llama 3.1 405B Instruct82.9
3Qwen 2.5 72B Instruct79.8
4Phi-3-small-8k-instruct79.1
5GPT-4.175.12
6Phi-3 Mini 4K Instruct71.7
7Llama 2 70B Base64.9
8GPT-3.5 Turbo61.59
9Phi-459.4
10Llama 2 7B58.5
11Mistral-7B-v0.156.1
12gemma-7B55.1
13Qwen 3 235B A22B55
14Yi 6B (Base)47.2
15falcon-180B37.1

Interactive version: theaggregate.ai/benchmark?slug=bbh · How the rankings work · Data refreshed daily, snapshot 2026-07-22.