XTC-Bench: leaderboard

Metric: Accuracy-weighted continuous cross-task agreement (AW-CCTA, 0-1, times 100): per-fact agreement between the model's generated image (scene graph extracted and judged) and its own answers to the matching questions, weighted by their joint correctness, averaged over all facts of XTC-Bench's 2,000 scene-graph-annotated images from COCO 2017 val and Visual Genome (over 31,000 atomic facts about objects, attributes and relations); a Qwen3-235B judge scores each fact 0-5, normalized to 0-1 and shown times 100; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 9 models tracked.

Top models

#ModelScore
1gemini-2.5-flash-image62.3
2BAGEL-7B-MoT56.9
3BLIP3-o8B42.3
4Janus-Pro-7B39
5OmniGen237.7
6Tar-7B35.8
7Show-o2-7B33.3
8MMaDA-8B14.4
9Show-o5.4

Interactive version: theaggregate.ai/benchmark?slug=xtc-bench · How It Works · Data refreshed daily, snapshot 2026-10-07.