CulMind - Seals: leaderboard

Metric: Seals subdomain score: macro-average of the task scores (0-100) in the subdomain, each of the 50 tasks keeping its own primary metric such as accuracy or F1, answer-only setting, images from more than 100 museum collections; higher is better. Source: arxiv.org. Saturation forecast: Around 2033. 14 models tracked.

Top models

#ModelScore
1Qwen 3 VL 32B Instruct53.1
2Gemini 3 Flash (Preview)51.5
3Qwen 3 VL 8B Instruct48.3
4GPT-5.546.5
5Gemini 2.5 Flash45.4
6GLM-4.5V45.2
7InternVL3-8B42.6
8InternVL3-14B41.6
9Qwen 3.5 27B41
10GPT-5.4 Mini39.7
11Qwen 3.5 9B37.1

Interactive version: theaggregate.ai/benchmark?slug=culmind-seals · How It Works · Data refreshed daily, snapshot 2026-09-29.