SADL: leaderboard

Metric: Average recall (%, AR): mean of the per-class recalls for distractor, excluded hard negative and non-distractor labels under guided classification (the model sees the image, the subject description and the enumerated candidates and labels each one), over the 14,617 annotated candidates of SADL (1,800 subject-aware cases on 1,000 photographs, including 1,938 hard negatives), zero-shot; higher is better. Source: arxiv.org. Saturation forecast: Around April 2027. 7 models tracked.

Top models

#ModelScore
1Gemini 3 Pro63.8
2GPT-4.162.3
3GPT-4o58
4Qwen 3 VL 235B A22B54
5Molmo2-8B46.9
6Llama 4 Scout43.7
7Pixtral-12B40.7

Interactive version: theaggregate.ai/benchmark?slug=sadl · How It Works · Data refreshed daily, snapshot 2026-09-29.