XTC-Bench - Object Retrieval: leaderboard

Metric: Understanding score (0-1, times 100) on object-retrieval questions (which object has the given attributes), a Qwen3-235B judge scores each fact 0-5, normalized to 0-1 and shown times 100, over XTC-Bench's 2,000 scene-graph-annotated images from COCO 2017 val and Visual Genome (over 31,000 atomic facts about objects, attributes and relations); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 10 models tracked.

Top models

#ModelScore
1GPT-571.8

Interactive version: theaggregate.ai/benchmark?slug=xtc-bench-object-retrieval · How It Works · Data refreshed daily, snapshot 2026-10-07.