SADL: leaderboard
Metric: Average recall (%, AR): mean of the per-class recalls for distractor, excluded hard negative and non-distractor labels under guided classification (the model sees the image, the subject description and the enumerated candidates and labels each one), over the 14,617 annotated candidates of SADL (1,800 subject-aware cases on 1,000 photographs, including 1,938 hard negatives), zero-shot; higher is better. Source: arxiv.org. Saturation forecast: Around April 2027. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Pro | 63.8 |
| 2 | GPT-4.1 | 62.3 |
| 3 | GPT-4o | 58 |
| 4 | Qwen 3 VL 235B A22B | 54 |
| 5 | Molmo2-8B | 46.9 |
| 6 | Llama 4 Scout | 43.7 |
| 7 | Pixtral-12B | 40.7 |
Interactive version: theaggregate.ai/benchmark?slug=sadl · How It Works · Data refreshed daily, snapshot 2026-09-29.