Appear2Meaning: leaderboard

Metric: Exact-match accuracy (%, times 100; every predicted metadata field judged correct) on the Appear2Meaning benchmark (750 Getty and Met museum objects, 50 per culture-type combination over four cultural regions and four object types; image-only input, structured JSON prediction; GPT-4.1-mini judge against the museum metadata); higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 9 models tracked.

Top models

#ModelScore
1Qwen 3 VL 32B Instruct2.9
2Qwen 3 VL 8B Instruct2.4
3GPT-4.1 Mini1.3
4Claude Haiku 4.51.2
5Pixtral-12B0.9
6GPT-5.4 Mini0.5

Interactive version: theaggregate.ai/benchmark?slug=appear2meaning · How It Works · Data refreshed daily, snapshot 2026-10-07.