XTC-Bench - Generation: leaderboard

Metric: Overall generation score (0-1, times 100): the model draws an image from a prompt verbalizing the reference scene graph, a scene graph is extracted from the image and matched to the reference, and a Qwen3-235B judge scores each fact 0-5, normalized to 0-1 and shown times 100 for every fact, over XTC-Bench's 2,000 scene-graph-annotated images from COCO 2017 val and Visual Genome (over 31,000 atomic facts about objects, attributes and relations); higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 10 models tracked.

Top models

#ModelScore
1GPT-Image-1.580.4
2gemini-2.5-flash-image79.5
3BAGEL-7B-MoT72.5
4BLIP3-o8B54.5
5Tar-7B51.8
6Janus-Pro-7B49.9
7OmniGen249.9
8Show-o2-7B46.5
9MMaDA-8B26.5
10Show-o6

Interactive version: theaggregate.ai/benchmark?slug=xtc-bench-generation · How It Works · Data refreshed daily, snapshot 2026-10-07.