360CityArena - Relational Spatial Reasoning: leaderboard

Metric: Accuracy (%; the agent names the landmark that stands in a given spatial relation to a reference landmark, judged semantically by GPT-5 and checked by hand; 25 human-authored tasks in a navigable reconstruction of the Akihabara district from 602 360-degree video segments; the agent acts with seven discrete view and movement actions and sees a map marking its position and heading; one run). Source: arxiv.org. Saturation forecast: Around March 2027. 6 models tracked.

Top models

#ModelScore
1GPT-532
2Gemini 2.5 Flash12
3Claude Sonnet 4.58
4Qwen 2.5 VL 32B Instruct4
5InternVL3.5-8B0

Interactive version: theaggregate.ai/benchmark?slug=360cityarena-relational-spatial-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-29.