CulMind - Ancient Books: leaderboard

Metric: Ancient Books subdomain score: macro-average of the task scores (0-100) in the subdomain, each of the 50 tasks keeping its own primary metric such as accuracy or F1, answer-only setting, images from more than 100 museum collections; higher is better. Source: arxiv.org. Saturation forecast: Around 2032. 14 models tracked.

Top models

#ModelScore
1Gemini 3 Flash (Preview)48.1
2Qwen 3 VL 32B Instruct45.8
3GPT-5.544.9
4GLM-4.5V44.3
5Gemini 2.5 Flash43.4
6Qwen 3 VL 8B Instruct43.4
7Qwen 3.5 27B39.6
8InternVL3-14B39.2
9GPT-5.4 Mini37.7
10InternVL3-8B37
11Qwen 3.5 9B33.1

Interactive version: theaggregate.ai/benchmark?slug=culmind-ancient-books · How It Works · Data refreshed daily, snapshot 2026-09-29.