XTC-Bench - Generation: leaderboard
Metric: Overall generation score (0-1, times 100): the model draws an image from a prompt verbalizing the reference scene graph, a scene graph is extracted from the image and matched to the reference, and a Qwen3-235B judge scores each fact 0-5, normalized to 0-1 and shown times 100 for every fact, over XTC-Bench's 2,000 scene-graph-annotated images from COCO 2017 val and Visual Genome (over 31,000 atomic facts about objects, attributes and relations); higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 10 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-Image-1.5 | 80.4 |
| 2 | gemini-2.5-flash-image | 79.5 |
| 3 | BAGEL-7B-MoT | 72.5 |
| 4 | BLIP3-o8B | 54.5 |
| 5 | Tar-7B | 51.8 |
| 6 | Janus-Pro-7B | 49.9 |
| 7 | OmniGen2 | 49.9 |
| 8 | Show-o2-7B | 46.5 |
| 9 | MMaDA-8B | 26.5 |
| 10 | Show-o | 6 |
Interactive version: theaggregate.ai/benchmark?slug=xtc-bench-generation · How It Works · Data refreshed daily, snapshot 2026-10-07.