BIG-Bench Extra Hard — leaderboard

BIG-Bench Extra Hard evaluates model capability on intelligence & reasoning tasks from the linked upstream source with Score as the primary reported metric.

Metric: Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 17 models tracked.

Top models

#ModelScore
1Gemma 4 31B74.4
2Gemma 4 26B A4B64.8
3Gemma 4 12B53
4O3 Mini (High)44.8
5Gemma 4 E4B33.1
6Gemma 4 E2B21.9
7Gemini 2.0 Flash9.8
8Gemini 2.0 Flash Lite8
9GPT-4o6
10Gemma 3 27B4.9
11Gemma 3 12B4.5
12Gemma 2 27B4
13Llama 3.1 8B Instruct3.6
14Gemma 3 4B3.4

Interactive version: theaggregate.ai/benchmark?slug=big-bench-extra-hard · How the rankings work · Data refreshed daily, snapshot 2026-07-22.