XTC-Bench - Object Retrieval: leaderboard
Metric: Understanding score (0-1, times 100) on object-retrieval questions (which object has the given attributes), a Qwen3-235B judge scores each fact 0-5, normalized to 0-1 and shown times 100, over XTC-Bench's 2,000 scene-graph-annotated images from COCO 2017 val and Visual Genome (over 31,000 atomic facts about objects, attributes and relations); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 10 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5 | 71.8 |
Interactive version: theaggregate.ai/benchmark?slug=xtc-bench-object-retrieval · How It Works · Data refreshed daily, snapshot 2026-10-07.