ZeroBench — leaderboard

"Impossible" visual benchmark for multimodal models. SOTA pass@5 is only 19%, 5/5 reliability just 6%. Designed to expose limits of current vision-language reasoning.

Metric: Score (%). Source: zerobench.github.io. Status: saturation imminent. 60 models tracked.

Top models

#ModelScore
1GPT-5.4 (xHigh)23
2Claude Fable 5 (Max)23
3GPT-5.5 (xHigh)22
4Gemini 3.1 Pro (Preview)19
5Gemini 3 Pro19
6Gemini 3.5 Flash19
7Claude Opus 4.817
8GPT-5.2 (Medium)17
9Claude Opus 4.7 (xHigh)14
10Gemini 3 Flash13
11Claude Sonnet 4.611
12Claude Opus 4.611
13Claude Opus 4.510
14GPT-5.4 Mini10
15GPT-5.1 (Medium)5

Interactive version: theaggregate.ai/benchmark?slug=zerobench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.