C3-Bench - Correctness: leaderboard
Metric: Correctness (how closely the caption matches the reference change) (0-10) of change captions on the 4,996 human-labelled C3-Bench image pairs (51 change contexts across natural scenes, remote sensing, image editing and anomalies), scored by a GPT-5.2 judge against the reference description with the context-specific change criteria; mean of three judge runs; higher is better. Source: arxiv.org. Saturation forecast: Around March 2028. 26 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.2 | 5.47 |
| 2 | Gemini 3 Pro | 5.45 |
| 3 | GPT-5.1 | 5.43 |
| 4 | Gemini 3 Flash | 5.42 |
| 5 | Gemini 2.5 Pro | 5.18 |
| 6 | Qwen 3 VL 32B Instruct | 5.18 |
| 7 | Qwen 3 VL 8B Instruct | 5.07 |
| 8 | Gemini 2.5 Flash | 4.99 |
| 9 | GPT-4o | 4.95 |
| 10 | GPT-5 Mini | 4.87 |
| 11 | InternVL3-78B | 4.86 |
| 12 | InternVL3-38B | 4.72 |
| 13 | Llama 4 Scout Instruct | 4.63 |
| 14 | GPT-4o Mini | 4.56 |
| 15 | Qwen 2.5 VL 72B Instruct | 4.54 |
Interactive version: theaggregate.ai/benchmark?slug=c3-bench-correctness · How It Works · Data refreshed daily, snapshot 2026-09-29.