ChronoVision: leaderboard

Metric: Overall score (0-100): mean of the Artifacts score (mean of dynasty accuracy and 50 x (1 + Kendall tau) ordering), the Shortcut score (color-pair accuracy x (1 - |grayscale accuracy gap|)) and the News score (mean of year accuracy discounted by MAE over 79 years and image-text matching accuracy); zero-shot; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 10 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro67.17
2GPT-5.249.96
3Qwen 3 VL 235B A22B Instruct49.92
4Qwen 3 VL 8B Instruct44.47
5Qwen 3 VL 4B Instruct41.19
6GLM-4.1V-9B (Thinking)37.35
7Qwen 2.5 VL 7B Instruct36.91
8InternVL3.5-8B29.06

Interactive version: theaggregate.ai/benchmark?slug=chronovision · How It Works · Data refreshed daily, snapshot 2026-09-29.