CulMind - Ancient Scripts: leaderboard
Metric: Ancient Scripts subdomain score: macro-average of the task scores (0-100) in the subdomain, each of the 50 tasks keeping its own primary metric such as accuracy or F1, answer-only setting, images from more than 100 museum collections; higher is better. Source: arxiv.org. Saturation forecast: Around 2043. 14 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 3 VL 8B Instruct | 19.1 |
| 2 | InternVL3-14B | 8.3 |
| 3 | InternVL3-8B | 7.8 |
| 4 | Gemini 3 Flash (Preview) | 7.1 |
| 5 | GPT-5.5 | 5.7 |
| 6 | Qwen 3 VL 32B Instruct | 5.1 |
| 7 | Gemini 2.5 Flash | 4.4 |
| 8 | Qwen 3.5 27B | 4 |
| 9 | GPT-5.4 Mini | 2.4 |
| 10 | GLM-4.5V | 1.4 |
| 11 | Qwen 3.5 9B | 0.8 |
Interactive version: theaggregate.ai/benchmark?slug=culmind-ancient-scripts · How It Works · Data refreshed daily, snapshot 2026-09-29.