Blind-Spots-Bench - Image Generation: leaderboard

Metric: Accuracy (%; mean@1 on the questions whose answer is a generated image, graded correct or incorrect by gemini-3-flash against the reference solution; one sample per question). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 6 models tracked.

Top models

#ModelScore
1gemini-3-pro-image-preview (nano-banana-pro)54.8
2gemini-3.1-flash-image-preview53.6
3GPT-image-251.2
4GPT-image-1.540.5
5gemini-2.5-flash-image22.6
6GPT-image-1-mini19

Interactive version: theaggregate.ai/benchmark?slug=blind-spots-bench-image-generation · How It Works · Data refreshed daily, snapshot 2026-09-29.