VIABench - Vision-Guided Interaction: leaderboard

Metric: Vision-guided interaction score (0-100): multi-turn guidance toward an intended interaction with the environment, with the ground-truth history of earlier turns supplied, judged by GPT-5 as the semantic similarity (0 to 1, reported times 100) between the model answer and the ground-truth annotation, on first-person videos recorded or shared by visually impaired people, with each model at its default frame sampling; higher is better. Source: arxiv.org. Saturation forecast: Around February 2027. 15 models tracked.

Top models

#ModelScore
1GPT-551.7
2Gemini 2.5 Pro43.5
3GPT-4o (2024-08-06)43.5
4Qwen 2.5 VL 7B27.1
5InternVL3.5-8B20.1

Interactive version: theaggregate.ai/benchmark?slug=viabench-vision-guided-interaction · How It Works · Data refreshed daily, snapshot 2026-09-29.