ChronoBench - Long-Term Memory: leaderboard
Metric: Accuracy (%; level average of object appear, object change and object history memory; 17,689 validated multiple-choice questions over remote-sensing image time series, box- and geo-coordinate-grounded variants; open models greedy with 128 new tokens, API models at high reasoning effort). Source: arxiv.org. Saturation forecast: Around March 2028. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Flash | 47.52 |
| 2 | GPT-5.4 (High) | 39.21 |
| 3 | Qwen 3 VL 32B Instruct | 24.06 |
| 4 | Qwen 3 VL 8B Instruct | 20.52 |
| 5 | Qwen 3 VL 4B Instruct | 17.54 |
| 6 | InternVL3.5-8B | 16.93 |
Interactive version: theaggregate.ai/benchmark?slug=chronobench-long-term-memory · How It Works · Data refreshed daily, snapshot 2026-09-29.