360CityArena - Relational Spatial Reasoning: leaderboard
Metric: Accuracy (%; the agent names the landmark that stands in a given spatial relation to a reference landmark, judged semantically by GPT-5 and checked by hand; 25 human-authored tasks in a navigable reconstruction of the Akihabara district from 602 360-degree video segments; the agent acts with seven discrete view and movement actions and sees a map marking its position and heading; one run). Source: arxiv.org. Saturation forecast: Around March 2027. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5 | 32 |
| 2 | Gemini 2.5 Flash | 12 |
| 3 | Claude Sonnet 4.5 | 8 |
| 4 | Qwen 2.5 VL 32B Instruct | 4 |
| 5 | InternVL3.5-8B | 0 |
Interactive version: theaggregate.ai/benchmark?slug=360cityarena-relational-spatial-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-29.