ArtECulture - Arabic: leaderboard

Metric: Accuracy (%; Arabic-speaking culture labels; zero-shot prediction of the majority emotion (nine categories) that the target culture perceives in each of the 679 test artworks; original models without the authors retrieval knowledge or fine-tuning). Source: arxiv.org. Saturation forecast: Around 2034. 16 models tracked.

Top models

#ModelScore
1Qwen 3.5 27B41.53
2Qwen 3.6 35B A3B39.76
3Gemini 3.5 Flash39.62
4Gemini 3 Flash39.18
5Gemma 4 31B (IT)38.44
6Qwen 3.6 27B38
7Claude Sonnet 4.635.94
8Claude Opus 4.833.04
9Qwen 3 VL 8B32.99
10GPT-5.531.96
11Qwen 3.5 9B31.22
12aya-vision-8B29.16
13Gemma 3 12B (IT)28.13
14Pixtral-12B27.69

Interactive version: theaggregate.ai/benchmark?slug=arteculture-arabic · How It Works · Data refreshed daily, snapshot 2026-09-26.