BBEH: leaderboard

BIG-Bench Extra Hard: 27 tasks from Google DeepMind that push frontier models beyond standard BIG-Bench Hard.

Metric: Harmonic Mean (%). Source: github.com. Status: saturated. 11 models tracked.

Top models

#ModelScore
1O3 Mini (High)44.8
2Gemini 2.0 Flash9.8
3Gemini 2.0 Flash Lite8
4DeepSeek R16.8
5GPT-4o6
6Gemma 3 27B4.9
7Gemma 3 12B4.5
8Gemma 2 27B (IT)4
9Llama 3.1 8B Instruct3.6
10Gemma 3 4B3.4

Interactive version: theaggregate.ai/benchmark?slug=bbeh · How It Works · Data refreshed daily, snapshot 2026-09-05.