XTC-Bench - Understanding: leaderboard
Metric: Overall understanding score (0-1, times 100) on visual questions derived one per fact from the reference scene graph of the real image, a Qwen3-235B judge scores each fact 0-5, normalized to 0-1 and shown times 100, over XTC-Bench's 2,000 scene-graph-annotated images from COCO 2017 val and Visual Genome (over 31,000 atomic facts about objects, attributes and relations); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 10 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5 | 68.4 |
Interactive version: theaggregate.ai/benchmark?slug=xtc-bench-understanding · How It Works · Data refreshed daily, snapshot 2026-10-07.