BBEH — leaderboard

BIG-Bench Extra Hard: 27 extremely challenging tasks from Google DeepMind designed to push frontier models beyond standard BIG-Bench Hard.

Metric: Harmonic Mean (%). Source: github.com. Status: saturation imminent. 11 models tracked.

Top models

#ModelScore
1O3 Mini (High)44.8
2Gemini 2.0 Flash9.8
3Gemini 2.0 Flash Lite8
4DeepSeek R16.8
5GPT-4o6
6Gemma 3 27B4.9
7Gemma 3 12B4.5
8Gemma 2 27B (IT)4
9Llama 3.1 8B Instruct3.6
10Gemma 3 4B3.4

Interactive version: theaggregate.ai/benchmark?slug=bbeh · How the rankings work · Data refreshed daily, snapshot 2026-07-22.