CulMind - Seals: leaderboard
Metric: Seals subdomain score: macro-average of the task scores (0-100) in the subdomain, each of the 50 tasks keeping its own primary metric such as accuracy or F1, answer-only setting, images from more than 100 museum collections; higher is better. Source: arxiv.org. Saturation forecast: Around 2033. 14 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 3 VL 32B Instruct | 53.1 |
| 2 | Gemini 3 Flash (Preview) | 51.5 |
| 3 | Qwen 3 VL 8B Instruct | 48.3 |
| 4 | GPT-5.5 | 46.5 |
| 5 | Gemini 2.5 Flash | 45.4 |
| 6 | GLM-4.5V | 45.2 |
| 7 | InternVL3-8B | 42.6 |
| 8 | InternVL3-14B | 41.6 |
| 9 | Qwen 3.5 27B | 41 |
| 10 | GPT-5.4 Mini | 39.7 |
| 11 | Qwen 3.5 9B | 37.1 |
Interactive version: theaggregate.ai/benchmark?slug=culmind-seals · How It Works · Data refreshed daily, snapshot 2026-09-29.