SPD-Faith Bench - Difference Sensitivity: leaderboard

Metric: Mean difference sensitivity (%), max(0, 1 - |reported count - true count| / true count) per pair, on the 1,000 multi-difference image pairs of SPD-Faith Bench (COCO 2017 images with two to five controlled edits each: colour change, object removal or position change, planned with an LLM, realized with LaMa inpainting and verified by human annotators); zero-shot chain-of-thought difference reasoning; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 12 models tracked.

Top models

#ModelScoreOverall rank
1GLM-4.5V92.5#339
2GPT-4o82#333
3Gemini 2.5 Pro79.7#145
4Claude Haiku 4.564.5#271
5Qwen 2.5 VL 72B Instruct44.4#364
6Qwen 2.5 VL 7B Instruct42.8#643
7MiniCPM-V-2.631.2#825

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=spd-faith-bench-difference-sensitivity · How It Works · Data refreshed daily, snapshot 2026-10-11.