ChronoVision: leaderboard
Metric: Overall score (0-100): mean of the Artifacts score (mean of dynasty accuracy and 50 x (1 + Kendall tau) ordering), the Shortcut score (color-pair accuracy x (1 - |grayscale accuracy gap|)) and the News score (mean of year accuracy discounted by MAE over 79 years and image-text matching accuracy); zero-shot; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 10 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Pro | 67.17 |
| 2 | GPT-5.2 | 49.96 |
| 3 | Qwen 3 VL 235B A22B Instruct | 49.92 |
| 4 | Qwen 3 VL 8B Instruct | 44.47 |
| 5 | Qwen 3 VL 4B Instruct | 41.19 |
| 6 | GLM-4.1V-9B (Thinking) | 37.35 |
| 7 | Qwen 2.5 VL 7B Instruct | 36.91 |
| 8 | InternVL3.5-8B | 29.06 |
Interactive version: theaggregate.ai/benchmark?slug=chronovision · How It Works · Data refreshed daily, snapshot 2026-09-29.