Blind-Spots-Bench - Image Generation: leaderboard
Metric: Accuracy (%; mean@1 on the questions whose answer is a generated image, graded correct or incorrect by gemini-3-flash against the reference solution; one sample per question). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | gemini-3-pro-image-preview (nano-banana-pro) | 54.8 |
| 2 | gemini-3.1-flash-image-preview | 53.6 |
| 3 | GPT-image-2 | 51.2 |
| 4 | GPT-image-1.5 | 40.5 |
| 5 | gemini-2.5-flash-image | 22.6 |
| 6 | GPT-image-1-mini | 19 |
Interactive version: theaggregate.ai/benchmark?slug=blind-spots-bench-image-generation · How It Works · Data refreshed daily, snapshot 2026-09-29.