ChronoBench - Temporal Recognition: leaderboard
Metric: Accuracy (%; level average of bi-temporal class-level change and area-change recognition, long-temporal class-level change recognition and bi-temporal object-level change recognition; 17,689 validated multiple-choice questions over remote-sensing image time series, box- and geo-coordinate-grounded variants; open models greedy with 128 new tokens, API models at high reasoning effort). Source: arxiv.org. Saturation forecast: Around June 2027. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.4 (High) | 67.57 |
| 2 | Gemini 3 Flash | 61.38 |
| 3 | Qwen 3 VL 32B Instruct | 57.88 |
| 4 | Qwen 3 VL 8B Instruct | 47.12 |
| 5 | InternVL3.5-8B | 43.21 |
| 6 | Qwen 3 VL 4B Instruct | 42.07 |
Interactive version: theaggregate.ai/benchmark?slug=chronobench-temporal-recognition · How It Works · Data refreshed daily, snapshot 2026-09-29.