Big-Bench Hard — leaderboard

23 challenging tasks from BIG-Bench where prior language models failed to outperform average human raters. Tests algorithmic reasoning, language understanding, and world knowledge.

Metric: Average (%). Source: epoch.ai. Status: saturation imminent. 50 models tracked.

Top models

#ModelScore
1Gemini 1.5 Pro (001)89.2
2DeepSeek V387.5
3Llama 3.1 405B82.9
4Phi-3-medium-128k-instruct81.4
5Qwen 2.5 72B79.8
6Phi-3-small-8k-instruct79.1
7DeepSeek V278.8
8GPT-4 (0613)75.12
9Phi-3 Mini 4K Instruct71.7
10Yi 34B (Chat)71.7
11StableBeluga269.3
12Llama 2 70B Base64.9
13GPT-3.5 Turbo (0613)61.59
14Phi-259.4
15Nemotron-4 15B58.7

Interactive version: theaggregate.ai/benchmark?slug=big-bench-hard · How the rankings work · Data refreshed daily, snapshot 2026-07-22.