C3-Bench - Relevance: leaderboard

Metric: Relevance (the caption covers the changes the context asks for and ignores the rest) (0-10) of change captions on the 4,996 human-labelled C3-Bench image pairs (51 change contexts across natural scenes, remote sensing, image editing and anomalies), scored by a GPT-5.2 judge against the reference description with the context-specific change criteria; mean of three judge runs; higher is better. Source: arxiv.org. Saturation forecast: Around October 2027. 26 models tracked.

Top models

#ModelScore
1GPT-5.25.98
2GPT-5.15.98
3Gemini 3 Pro5.95
4Gemini 3 Flash5.86
5Gemini 2.5 Pro5.85
6GPT-4o5.83
7Gemini 2.5 Flash5.69
8InternVL3-78B5.59
9Qwen 3 VL 32B Instruct5.58
10Qwen 3 VL 8B Instruct5.54
11InternVL3-38B5.46
12GPT-4o Mini5.23
13Llama 4 Scout Instruct5.19
14Qwen 2.5 VL 7B Instruct5.07
15Qwen 2.5 VL 72B Instruct5.06

Interactive version: theaggregate.ai/benchmark?slug=c3-bench-relevance · How It Works · Data refreshed daily, snapshot 2026-09-29.