C3-Bench - Relevance: leaderboard
Metric: Relevance (the caption covers the changes the context asks for and ignores the rest) (0-10) of change captions on the 4,996 human-labelled C3-Bench image pairs (51 change contexts across natural scenes, remote sensing, image editing and anomalies), scored by a GPT-5.2 judge against the reference description with the context-specific change criteria; mean of three judge runs; higher is better. Source: arxiv.org. Saturation forecast: Around October 2027. 26 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.2 | 5.98 |
| 2 | GPT-5.1 | 5.98 |
| 3 | Gemini 3 Pro | 5.95 |
| 4 | Gemini 3 Flash | 5.86 |
| 5 | Gemini 2.5 Pro | 5.85 |
| 6 | GPT-4o | 5.83 |
| 7 | Gemini 2.5 Flash | 5.69 |
| 8 | InternVL3-78B | 5.59 |
| 9 | Qwen 3 VL 32B Instruct | 5.58 |
| 10 | Qwen 3 VL 8B Instruct | 5.54 |
| 11 | InternVL3-38B | 5.46 |
| 12 | GPT-4o Mini | 5.23 |
| 13 | Llama 4 Scout Instruct | 5.19 |
| 14 | Qwen 2.5 VL 7B Instruct | 5.07 |
| 15 | Qwen 2.5 VL 72B Instruct | 5.06 |
Interactive version: theaggregate.ai/benchmark?slug=c3-bench-relevance · How It Works · Data refreshed daily, snapshot 2026-09-29.