SPD-Faith Bench - Type-Level F1: leaderboard

Metric: Type-level F1 (%), the unweighted mean of the colour-change, object-removal and position-change F1 scores (the printed Overall column; the text calls it a micro average, but every row's Overall precision, recall and F1 are the means of the three type columns), on the 1,000 multi-difference image pairs of SPD-Faith Bench (COCO 2017 images with two to five controlled edits each: colour change, object removal or position change, planned with an LLM, realized with LaMa inpainting and verified by human annotators); zero-shot chain-of-thought difference reasoning; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 12 models tracked.

Top models

#ModelScoreOverall rank
1GLM-4.5V69.4#339
2Gemini 2.5 Pro66.1#145
3GPT-4o59.7#333
4Qwen 2.5 VL 72B Instruct56.6#364
5Claude Haiku 4.554.7#271
6Qwen 2.5 VL 7B Instruct51.6#643
7MiniCPM-V-2.639.6#825

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=spd-faith-bench-type-level-f1 · How It Works · Data refreshed daily, snapshot 2026-10-11.