C3-Bench - Specificity: leaderboard

Metric: Specificity (a concrete explanation rather than a generic statement) (0-10) of change captions on the 4,996 human-labelled C3-Bench image pairs (51 change contexts across natural scenes, remote sensing, image editing and anomalies), scored by a GPT-5.2 judge against the reference description with the context-specific change criteria; mean of three judge runs; higher is better. Source: arxiv.org. Saturation forecast: Around 2028. 26 models tracked.

Top models

#ModelScore
1GPT-5.25.16
2Gemini 3 Pro5.15
3GPT-5.15.12
4Gemini 3 Flash5.08
5Qwen 3 VL 32B Instruct4.93
6Gemini 2.5 Pro4.92
7Qwen 3 VL 8B Instruct4.76
8Gemini 2.5 Flash4.69
9GPT-5 Mini4.6
10InternVL3-78B4.51
11Llama 4 Scout Instruct4.43
12Qwen 2.5 VL 72B Instruct4.38
13GPT-4o4.34
14InternVL3-38B4.34
15Qwen 2.5 VL 32B Instruct4.28

Interactive version: theaggregate.ai/benchmark?slug=c3-bench-specificity · How It Works · Data refreshed daily, snapshot 2026-09-29.