ArtECulture - English: leaderboard

Metric: Accuracy (%; English-speaking culture labels; zero-shot prediction of the majority emotion (nine categories) that the target culture perceives in each of the 679 test artworks; original models without the authors retrieval knowledge or fine-tuning). Source: arxiv.org. Saturation forecast: Around 2031. 16 models tracked.

Top models

#ModelScore
1Gemma 4 31B (IT)59.35
2Claude Opus 4.859.2
3Gemini 3.5 Flash59.2
4Claude Sonnet 4.658.62
5GPT-5.558.62
6Gemini 3 Flash58.62
7Qwen 3.5 27B58.32
8Qwen 3.6 27B57.58
9Qwen 3.6 35B A3B55.38
10Qwen 3.5 9B55.23
11aya-vision-8B54.34
12Gemma 3 12B (IT)53.76
13Pixtral-12B52.72
14Qwen 3 VL 8B50.37

Interactive version: theaggregate.ai/benchmark?slug=arteculture-english · How It Works · Data refreshed daily, snapshot 2026-09-26.