LAD-Bench - Unprompted Detection: leaderboard

Metric: Level 1 accuracy (%): share of images whose anomaly the model points out when given the image alone, with no text prompt or hint; LAD-Bench: 1,007 synthetic images (gpt-image-1) with one logical anomaly each (nature, residential, urban and collaborative scenes); free-text answers graded by GPT-5-nano against the annotated anomaly; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 7 models tracked.

Top models

#ModelScore
1Gemini 3 Flash (Preview)54.22
2Claude Sonnet 4.645.58
3GPT-544.79
4GPT-5 Mini35.05
5Grok 4.1 Fast (Reasoning)28.6

Interactive version: theaggregate.ai/benchmark?slug=lad-bench-unprompted-detection · How It Works · Data refreshed daily, snapshot 2026-09-29.