IntentionNav: leaderboard

Metric: Success rate (%): the agent stops within 2.0 m of the scene-grounded target, over IntentionNav's 2,000 episodes (500 implicit-intent instructions in four styles, 176 Isaac Sim scenes, 64 target categories, target name withheld), the VLM driving the benchmark's fixed reference active-navigation agent (RGB-D, open-vocabulary detection, map memory, 30-step budget); higher is better. Source: arxiv.org. Saturation forecast: Around June 2028. 3 models tracked.

Top models

#ModelScore
1Gemini 3.1 Flash Lite25.7
2GPT-5.424.9
3Qwen 3.6 Plus24.1

Interactive version: theaggregate.ai/benchmark?slug=intentionnav · How It Works · Data refreshed daily, snapshot 2026-10-07.