LAD-Bench: leaderboard
Metric: Decay-weighted accuracy (%): an anomaly found with no hint scores 100, after being told the image has an anomaly 66.7, after a category hint 33.3, otherwise 0; LAD-Bench: 1,007 synthetic images (gpt-image-1) with one logical anomaly each (nature, residential, urban and collaborative scenes); free-text answers graded by GPT-5-nano against the annotated anomaly; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 4.6 | 70.11 |
| 2 | Gemini 3 Flash (Preview) | 67.33 |
| 3 | GPT-5 | 67.2 |
| 4 | GPT-5 Mini | 65.96 |
| 5 | Grok 4.1 Fast (Reasoning) | 55.28 |
Interactive version: theaggregate.ai/benchmark?slug=lad-bench · How It Works · Data refreshed daily, snapshot 2026-09-29.