CArtBench - CuratorQA (Subject Recognition): leaderboard

Metric: Exact-match accuracy (%) on the subject and object recognition (QA1) CuratorQA questions of CArtBench (Palace Museum artworks with catalog-grounded multiple-choice and true-or-false questions, answered from the artwork image, constrained decoding, temperature 0); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 9 models tracked.

Top models

#ModelScore
1Qwen 3 VL 235B A22B (Thinking)88
2GPT-5 Mini84
3Qwen 2.5 VL 7B82
4GPT-5 Nano62
5GLM-4.5V59
6Gemini 2.5 Flash51

Interactive version: theaggregate.ai/benchmark?slug=cartbench-curatorqa-subject-recognition · How It Works · Data refreshed daily, snapshot 2026-10-07.