ChronoBench - Temporal Recognition: leaderboard

Metric: Accuracy (%; level average of bi-temporal class-level change and area-change recognition, long-temporal class-level change recognition and bi-temporal object-level change recognition; 17,689 validated multiple-choice questions over remote-sensing image time series, box- and geo-coordinate-grounded variants; open models greedy with 128 new tokens, API models at high reasoning effort). Source: arxiv.org. Saturation forecast: Around June 2027. 11 models tracked.

Top models

#ModelScore
1GPT-5.4 (High)67.57
2Gemini 3 Flash61.38
3Qwen 3 VL 32B Instruct57.88
4Qwen 3 VL 8B Instruct47.12
5InternVL3.5-8B43.21
6Qwen 3 VL 4B Instruct42.07

Interactive version: theaggregate.ai/benchmark?slug=chronobench-temporal-recognition · How It Works · Data refreshed daily, snapshot 2026-09-29.