VIABench - Vision-Guided Interaction: leaderboard
Metric: Vision-guided interaction score (0-100): multi-turn guidance toward an intended interaction with the environment, with the ground-truth history of earlier turns supplied, judged by GPT-5 as the semantic similarity (0 to 1, reported times 100) between the model answer and the ground-truth annotation, on first-person videos recorded or shared by visually impaired people, with each model at its default frame sampling; higher is better. Source: arxiv.org. Saturation forecast: Around February 2027. 15 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5 | 51.7 |
| 2 | Gemini 2.5 Pro | 43.5 |
| 3 | GPT-4o (2024-08-06) | 43.5 |
| 4 | Qwen 2.5 VL 7B | 27.1 |
| 5 | InternVL3.5-8B | 20.1 |
Interactive version: theaggregate.ai/benchmark?slug=viabench-vision-guided-interaction · How It Works · Data refreshed daily, snapshot 2026-09-29.