C3-Bench - Specificity: leaderboard
Metric: Specificity (a concrete explanation rather than a generic statement) (0-10) of change captions on the 4,996 human-labelled C3-Bench image pairs (51 change contexts across natural scenes, remote sensing, image editing and anomalies), scored by a GPT-5.2 judge against the reference description with the context-specific change criteria; mean of three judge runs; higher is better. Source: arxiv.org. Saturation forecast: Around 2028. 26 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.2 | 5.16 |
| 2 | Gemini 3 Pro | 5.15 |
| 3 | GPT-5.1 | 5.12 |
| 4 | Gemini 3 Flash | 5.08 |
| 5 | Qwen 3 VL 32B Instruct | 4.93 |
| 6 | Gemini 2.5 Pro | 4.92 |
| 7 | Qwen 3 VL 8B Instruct | 4.76 |
| 8 | Gemini 2.5 Flash | 4.69 |
| 9 | GPT-5 Mini | 4.6 |
| 10 | InternVL3-78B | 4.51 |
| 11 | Llama 4 Scout Instruct | 4.43 |
| 12 | Qwen 2.5 VL 72B Instruct | 4.38 |
| 13 | GPT-4o | 4.34 |
| 14 | InternVL3-38B | 4.34 |
| 15 | Qwen 2.5 VL 32B Instruct | 4.28 |
Interactive version: theaggregate.ai/benchmark?slug=c3-bench-specificity · How It Works · Data refreshed daily, snapshot 2026-09-29.