C3-Bench - Correctness: leaderboard

Metric: Correctness (how closely the caption matches the reference change) (0-10) of change captions on the 4,996 human-labelled C3-Bench image pairs (51 change contexts across natural scenes, remote sensing, image editing and anomalies), scored by a GPT-5.2 judge against the reference description with the context-specific change criteria; mean of three judge runs; higher is better. Source: arxiv.org. Saturation forecast: Around March 2028. 26 models tracked.

Top models

#ModelScore
1GPT-5.25.47
2Gemini 3 Pro5.45
3GPT-5.15.43
4Gemini 3 Flash5.42
5Gemini 2.5 Pro5.18
6Qwen 3 VL 32B Instruct5.18
7Qwen 3 VL 8B Instruct5.07
8Gemini 2.5 Flash4.99
9GPT-4o4.95
10GPT-5 Mini4.87
11InternVL3-78B4.86
12InternVL3-38B4.72
13Llama 4 Scout Instruct4.63
14GPT-4o Mini4.56
15Qwen 2.5 VL 72B Instruct4.54

Interactive version: theaggregate.ai/benchmark?slug=c3-bench-correctness · How It Works · Data refreshed daily, snapshot 2026-09-29.