ArtECulture - Chinese: leaderboard

Metric: Accuracy (%; Chinese-speaking culture labels; zero-shot prediction of the majority emotion (nine categories) that the target culture perceives in each of the 679 test artworks; original models without the authors retrieval knowledge or fine-tuning). Source: arxiv.org. Saturation forecast: Around 2030. 16 models tracked.

Top models

#ModelScore
1Claude Opus 4.851.33
2Gemini 3 Flash51.25
3Gemini 3.5 Flash50.81
4Gemma 4 31B (IT)46.24
5Qwen 3.5 27B45.36
6Claude Sonnet 4.645.21
7Qwen 3.5 9B43.89
8Qwen 3.6 35B A3B43
9GPT-5.541.83
10Qwen 3 VL 8B41.38
11Qwen 3.6 27B38.73
12Gemma 3 12B (IT)36.23
13aya-vision-8B33.87
14Pixtral-12B17.08

Interactive version: theaggregate.ai/benchmark?slug=arteculture-chinese · How It Works · Data refreshed daily, snapshot 2026-09-26.