SPD-Faith Bench - Category-Level F1: leaderboard
Metric: Micro F1 (%) of the COCO object categories the model reports as modified against the edited objects, on the 1,000 multi-difference image pairs of SPD-Faith Bench (COCO 2017 images with two to five controlled edits each: colour change, object removal or position change, planned with an LLM, realized with LaMa inpainting and verified by human annotators); zero-shot chain-of-thought difference reasoning; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 11 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | GLM-4.5V | 69.6 | #339 |
| 2 | Gemini 2.5 Pro | 62.8 | #145 |
| 3 | GPT-4o | 55.6 | #333 |
| 4 | Qwen 2.5 VL 72B Instruct | 55.5 | #364 |
| 5 | Claude Haiku 4.5 | 51.4 | #271 |
| 6 | Qwen 2.5 VL 7B Instruct | 38.3 | #643 |
| 7 | MiniCPM-V-2.6 | 26.1 | #825 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=spd-faith-bench-category-level-f1 · How It Works · Data refreshed daily, snapshot 2026-10-11.