IntentionNav: leaderboard
Metric: Success rate (%): the agent stops within 2.0 m of the scene-grounded target, over IntentionNav's 2,000 episodes (500 implicit-intent instructions in four styles, 176 Isaac Sim scenes, 64 target categories, target name withheld), the VLM driving the benchmark's fixed reference active-navigation agent (RGB-D, open-vocabulary detection, map memory, 30-step budget); higher is better. Source: arxiv.org. Saturation forecast: Around June 2028. 3 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.1 Flash Lite | 25.7 |
| 2 | GPT-5.4 | 24.9 |
| 3 | Qwen 3.6 Plus | 24.1 |
Interactive version: theaggregate.ai/benchmark?slug=intentionnav · How It Works · Data refreshed daily, snapshot 2026-10-07.