360CityArena - Vision-Language Navigation: leaderboard

Metric: Success rate (%; the agent must follow multi-step natural-language route instructions and stop within 10 m of the goal; 25 human-authored tasks in a navigable reconstruction of the Akihabara district from 602 360-degree video segments; the agent acts with seven discrete view and movement actions and sees a map marking its position and heading; one run). Source: arxiv.org. Saturation forecast: Around 2033. 6 models tracked.

Top models

#ModelScore
1InternVL3.5-8B12
2GPT-58
3Gemini 2.5 Flash8
4Claude Sonnet 4.54
5Qwen 2.5 VL 32B Instruct0

Interactive version: theaggregate.ai/benchmark?slug=360cityarena-vision-language-navigation · How It Works · Data refreshed daily, snapshot 2026-09-29.