Pix2Fact (No Search): leaderboard

Metric: Accuracy (%) on the 1,000 Pix2Fact questions (one per 4K-to-8K real-world photo across eight urban scene categories; each question needs a pixel-level visual clue and external knowledge, written and reviewed by PhD annotators), answers judged against the reference by a Gemini-3.1-Pro LLM judge; temperature 0 (GPT-5.4 at its fixed temperature 1), images resized to API or 15-megapixel limits where needed, and an API call failing three times counts as wrong; original full image, no web search (condition C1); higher is better. Source: arxiv.org. Saturation forecast: Around May 2028. 10 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3.1 Pro (Preview)18.4#54
2Gemini 2.5 Pro14.6#145
3Claude Opus 4.713.3#45
4GPT-5.48.5#76
5Seed 2.0 Pro8#111
6Qwen 3.6 27B4.7#163
7Grok 4.204.4#167
8GLM-4.6V2.9#309
9Gemma 4 31B (IT)2.8#171

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=pix2fact-no-search · How It Works · Data refreshed daily, snapshot 2026-10-11.