ChronoBench - Spatio-Temporal Reasoning: leaderboard

Metric: Accuracy (%; level average of class-level change magnitude estimation, region development comparison and object construction ordering within and across sequences; 17,689 validated multiple-choice questions over remote-sensing image time series, box- and geo-coordinate-grounded variants; open models greedy with 128 new tokens, API models at high reasoning effort). Source: arxiv.org. Saturation forecast: Around 2029. 11 models tracked.

Top models

#ModelScore
1Gemini 3 Flash59.89
2GPT-5.4 (High)50.42
3Qwen 3 VL 32B Instruct47.23
4InternVL3.5-8B45.11
5Qwen 3 VL 8B Instruct42.8
6Qwen 3 VL 4B Instruct39.53

Interactive version: theaggregate.ai/benchmark?slug=chronobench-spatio-temporal-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-29.