LAD-Bench: leaderboard

Metric: Decay-weighted accuracy (%): an anomaly found with no hint scores 100, after being told the image has an anomaly 66.7, after a category hint 33.3, otherwise 0; LAD-Bench: 1,007 synthetic images (gpt-image-1) with one logical anomaly each (nature, residential, urban and collaborative scenes); free-text answers graded by GPT-5-nano against the annotated anomaly; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 7 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.670.11
2Gemini 3 Flash (Preview)67.33
3GPT-567.2
4GPT-5 Mini65.96
5Grok 4.1 Fast (Reasoning)55.28

Interactive version: theaggregate.ai/benchmark?slug=lad-bench · How It Works · Data refreshed daily, snapshot 2026-09-29.