LAD-Bench - Unprompted Detection: leaderboard
Metric: Level 1 accuracy (%): share of images whose anomaly the model points out when given the image alone, with no text prompt or hint; LAD-Bench: 1,007 synthetic images (gpt-image-1) with one logical anomaly each (nature, residential, urban and collaborative scenes); free-text answers graded by GPT-5-nano against the annotated anomaly; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Flash (Preview) | 54.22 |
| 2 | Claude Sonnet 4.6 | 45.58 |
| 3 | GPT-5 | 44.79 |
| 4 | GPT-5 Mini | 35.05 |
| 5 | Grok 4.1 Fast (Reasoning) | 28.6 |
Interactive version: theaggregate.ai/benchmark?slug=lad-bench-unprompted-detection · How It Works · Data refreshed daily, snapshot 2026-09-29.