ChronoBench: leaderboard
Metric: Overall accuracy (%; all 17,689 validated multiple-choice questions over remote-sensing image time series, box- and geo-coordinate-grounded variants; open models greedy with 128 new tokens, API models at high reasoning effort). Source: arxiv.org. Saturation forecast: Around September 2028. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.4 (High) | 56.29 |
| 2 | Qwen 3 VL 32B Instruct | 46.73 |
| 3 | Qwen 3 VL 8B Instruct | 39.19 |
| 4 | InternVL3.5-8B | 36.32 |
| 5 | Qwen 3 VL 4B Instruct | 35.1 |
Interactive version: theaggregate.ai/benchmark?slug=chronobench · How It Works · Data refreshed daily, snapshot 2026-09-29.