ArtECulture: leaderboard

Metric: Accuracy (%; mean of the English, Chinese and Arabic culture accuracies; zero-shot prediction of the majority emotion (nine categories) that the target culture perceives in each of the 679 test artworks; original models without the authors retrieval knowledge or fine-tuning). Source: arxiv.org. Saturation forecast: Around 2033. 16 models tracked.

Top models

#ModelScore
1Gemini 3.5 Flash49.88
2Gemini 3 Flash49.68
3Qwen 3.5 27B48.4
4Gemma 4 31B (IT)48.01
5Claude Opus 4.847.86
6Claude Sonnet 4.646.59
7Qwen 3.6 35B A3B46.05
8Qwen 3.6 27B44.77
9GPT-5.544.13
10Qwen 3.5 9B43.45
11Qwen 3 VL 8B41.58
12Gemma 3 12B (IT)39.37
13aya-vision-8B39.13
14Pixtral-12B32.5

Interactive version: theaggregate.ai/benchmark?slug=arteculture · How It Works · Data refreshed daily, snapshot 2026-09-26.