ChronoBench - Long-Term Memory: leaderboard

Metric: Accuracy (%; level average of object appear, object change and object history memory; 17,689 validated multiple-choice questions over remote-sensing image time series, box- and geo-coordinate-grounded variants; open models greedy with 128 new tokens, API models at high reasoning effort). Source: arxiv.org. Saturation forecast: Around March 2028. 11 models tracked.

Top models

#ModelScore
1Gemini 3 Flash47.52
2GPT-5.4 (High)39.21
3Qwen 3 VL 32B Instruct24.06
4Qwen 3 VL 8B Instruct20.52
5Qwen 3 VL 4B Instruct17.54
6InternVL3.5-8B16.93

Interactive version: theaggregate.ai/benchmark?slug=chronobench-long-term-memory · How It Works · Data refreshed daily, snapshot 2026-09-29.